zoukankan html css js c++ java

Scrapy提取多个标签的text

对于要提取嵌套标签所有内容的情况, 使用string或//text(), 注意两者区别

>>> from scrapy import Selector
>>> 
>>> doc = "<p id='test'>hello<b>world!</b></p>"
>>> 
>>> sel = Selector(text=doc, type='html')
>>> 
>>> sel.xpath("/p[@id='test']/text()").extract()
[]

使用text()

>>>#使用两个反斜杠
>>> sel.xpath("//p[@id='test']/text()").extract()
[u'hello']
>>> #这样提取出来是一个列表, 
>>> sel.xpath("//p[@id='test']//text()").extract()
[u'hello', u'world!']
>>>

使用string

>>> sel.xpath("//p[@id='test']").xpath('string(.)').extract()
[u'helloworld!']
>>> 
>>> sel.xpath("string(//p[@id='test'])").extract()
[u'helloworld!']
>>>

查看全文

相关阅读:
Jenkins Install
提高C#代码质量的22条准则
 游戏程序员英文指南
 苹果设备内存指南
 Unity符号表
 UI优化策略-UI性能优化技巧
 C# 语言历史版本特性
 CPU SIMD介绍
 Unity渲染性能指标
 关于JMeter线程组中线程数，Ramp-Up Period，循环次数之间的设置概念

原文地址：https://www.cnblogs.com/qlshine/p/5926101.html