zoukankan      html  css  js  c++  java
  • 通过脚本同时运行几个spider

    # 通过脚本同时运行几个spider
    目录结构:

    1.在命令行能通过的情况下创建两个spider如
    TestSpider
    Test2Spider


    2.在items.py的同级目录创建run.py文件,有三种方式,任选其一,其代码如下:

    方式一: 通过CrawlerProcess同时运行几个spider

    run_by_CrawlerProcess.py源代码:

     1 # 通过CrawlerProcess同时运行几个spider
     2 from scrapy.crawler import CrawlerProcess
     3 # 导入获取项目配置的模块
     4 from scrapy.utils.project import get_project_settings
     5 # 导入蜘蛛模块(即自己创建的spider)
     6 from spiders.test import TestSpider
     7 from spiders.test2 import Test2Spider
     8 
     9 # get_project_settings() 必须得有,不然"HTTP status code is not handled or not allowed"
    10 process = CrawlerProcess(get_project_settings())
    11 process.crawl(TestSpider) # 注意引入
    12 #process.crawl(Test2Spider) # 注意引入
    13 process.start()

    方式二:通过CrawlerRunner同时运行几个spider

     run_by_CrawlerRunner.py源代码:

     1 # 通过CrawlerRunner同时运行几个spider
     2 from twisted.internet import reactor
     3 from scrapy.crawler import CrawlerRunner
     4 from scrapy.utils.log import configure_logging
     5 # 导入获取项目配置的模块
     6 from scrapy.utils.project import get_project_settings
     7 # 导入蜘蛛模块(即自己创建的spider)
     8 from spiders.test import TestSpider
     9 from spiders.test2 import Test2Spider
    10 
    11 configure_logging()
    12 # get_project_settings() 必须得有,不然"HTTP status code is not handled or not allowed"
    13 runner = CrawlerRunner(get_project_settings())
    14 runner.crawl(TestSpider)
    15 #runner.crawl(Test2Spider)
    16 d = runner.join()
    17 d.addBoth(lambda _: reactor.stop())
    18 reactor.run() # the script will block here until all crawling jobs are finished

    方式三:通过CrawlerRunner和链接(chaining) deferred来线性运行来同时运行几个spider

    run_by_CrawlerRunner_and_Deferred.py源代码:

     1 # 通过CrawlerRunner和链接(chaining) deferred来线性运行来同时运行几个spider
     2 from twisted.internet import reactor, defer
     3 from scrapy.crawler import CrawlerRunner
     4 from scrapy.utils.log import configure_logging
     5 # 导入获取项目配置的模块
     6 from scrapy.utils.project import get_project_settings
     7 # 导入蜘蛛模块(即自己创建的spider)
     8 from spiders.test import TestSpider
     9 from spiders.test2 import Test2Spider
    10 
    11 configure_logging()
    12 # get_project_settings() 必须得有,不然"HTTP status code is not handled or not allowed"
    13 runner = CrawlerRunner(get_project_settings())
    14 
    15 @defer.inlineCallbacks
    16 def crawl():
    17     yield runner.crawl(TestSpider) 
    18     #yield runner.crawl(Test2Spider) 
    19     reactor.stop()
    20 
    21 crawl()
    22 reactor.run() # the script will block here until the last crawl call is finished

    3.修改两个spider文件引入items,和外部类的如(HeadersHelper.py)的引入模式(以run.py所在目录为中心)
    原导入模式:

    from ..items import ScrapydoubanmovieItem
    from .HeadersHelper import HeadersHelper

    注释:这种导入能够在命令行scrapy crawl Test正常运行

    修改为:

    from items import ScrapydoubanmovieItem
    from .HeadersHelper import HeadersHelper

    注释:修改后这种导入在命令行scrapy crawl Test会报错,但通过运行run.py文件,能够同时运行两个spider


    4.按照运行python文件的方式运行run.py,可以得到结果

  • 相关阅读:
    Log4Net使用指南
    Log4net 写文件日志与数据库日志
    JSON-Schema 最科学的表单验证模式
    番茄时间工作法
    css 温故而知新 1px的问题
    $.ajax 温故而知新坑
    H5中滚动卡顿的问题
    横向思维
    Wd 西部数据
    使用AlloyLever来搞定开发调试发布,错误监控上报,用户问题定位
  • 原文地址:https://www.cnblogs.com/xiaomingzaixian/p/7277615.html
Copyright © 2011-2022 走看看