Generator: 0 records selected for fetching, exiting ... Stopping at depth=0 - no more URLs to fetch.
出现上面的错误一般都会是nutch/conf/crawl-urlfilter.txt中的配置出现的不可预见的错误
我在网上找了好多配置发现
# accept hosts in MY.DOMAIN.NAME +^http://([a-z0-9]*/.)*360buy.com/
([a-z0-9]*/.)里的/这个写错了,正确的如下:
# accept hosts in MY.DOMAIN.NAME +^http://([a-z0-9]*.)*qq.com/