zoukankan html css js c++ java

fake-useragent User Agent 伪装

前几天意外找到一个简单实用的库- fake-useragent,可以伪装生成headers请求头中的User Agent值。

安装

pip3 install fake-useragent

各浏览器的user-agent值

from fake_useragent import UserAgent
ua = UserAgent()

ie浏览器的user agent

print(ua.ie)
Mozilla/5.0 (Windows; U; MSIE 9.0; Windows NT 9.0; en-US)

opera浏览器

print(ua.opera)
Opera/9.80 (X11; Linux i686; U; ru) Presto/2.8.131 Version/11.11

chrome浏览器

print(ua.chrome)
Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.2 (KHTML, like Gecko) Chrome/22.0.1216.0 Safari/537.2

firefox浏览器

print(ua.firefox)
Mozilla/5.0 (Windows NT 6.2; Win64; x64; rv:16.0.1) Gecko/20121011 Firefox/16.0.1

safri浏览器

print(ua.safari)
Mozilla/5.0 (iPad; CPU OS 6_0 like Mac OS X) AppleWebKit/536.26 (KHTML, like Gecko) Version/6.0 Mobile/10A5355d Safari/8536.25

最实用的

但我认为写爬虫最实用的是可以随意变换headers，一定要有随机性。在这里我写了三个随机生成user agent，三次打印都不一样，随机性很强，十分方便。

from fake_useragent import UserAgent
ua = UserAgent()
print(ua.random)
print(ua.random)
print(ua.random)
Mozilla/5.0 (X11; Ubuntu; Linux i686; rv:15.0) Gecko/20100101 Firefox/15.0.1
Mozilla/5.0 (Windows NT 6.2; Win64; x64; rv:16.0.1) Gecko/20121011 Firefox/16.0.1
Opera/9.80 (X11; Linux i686; U; ru) Presto/2.8.131 Version/11.11

爬虫中具体使用方法

import requests
from fake_useragent import UserAgent
ua = UserAgent()
headers = {'User-Agent': ua.random}
url = '待爬网页的url'
resp = requests.get(url, headers=headers)

省略具体爬虫的解析代码，大家可以回去试试

参考链接：点我

查看全文

相关阅读:
sql语句性能优化
 Windows版Redis如何使用？（单机）
redis在项目中的使用（单机版、集群版）
在windows上搭建redis集群(redis-cluster)
Jenkins打包Maven项目
 numpy交换列
 Linq中join多字段匹配
 SpringMVC Web项目升级为Springboot项目（二）
SpringMVC Web项目升级为Springboot项目（一）
springboot读取application.properties中自定义配置

原文地址：https://www.cnblogs.com/zswbky/p/8454085.html