zoukankan      html  css  js  c++  java
  • 2019-02-01 Python爬虫爬取豆瓣Top250

    这几天学了一点爬虫后写了个爬取电影top250的代码,分别用requests库和urllib库,想看看自己能不能搞出个啥东西,虽然很简单但还是小开心。

    import requests
    import re
    
    # https://movie.douban.com/top250?start=25&filter=
    # <span class="title">勇士</span>
    
    count = 1
    
    
    def getdata(url):
        data = requests.get(url)
        return data.text
    
    
    def showdata(data):
        global count
        regex = re.compile(r"<span class="title">(.*?)</span>")
        data = regex.findall(data)
        newdata = data.copy()
        for dataa in newdata:
            if "nbsp" in dataa:
                data.remove(dataa)
        for i in data:
            print(count, i)
            count = count + 1
    
    
    for i in range(0, 10):
        i = i * 25
        url = "https://movie.douban.com/top250?start={}&filter=".format(str(i))
        data = getdata(url)
        showdata(data)
    
    # 用requests来实现,正则表达式解析网页
    
    import urllib
    import urllib.request
    import re
    #https://movie.douban.com/top250?start=25&filter=
    #<span class="title">勇士</span>
    
    count = 1
    def getdata(url):
        data = urllib.request.urlopen(url).read().decode("utf-8")
        return data
    
    
    def showdata(data):
        global count
        regex = re.compile(r"<span class="title">(.*?)</span>")
        data = regex.findall(data)
        newdata = data.copy()
        for dataa in newdata:
            if "nbsp" in dataa:
                data.remove(dataa)
        for i in data:
            print(count,i)
            count = count+1
    
    
    for i in range(0,10):
        i = i*25
        url = "https://movie.douban.com/top250?start={}&filter=".format(str(i))
        data = getdata(url)
        showdata(data)
    
    #用urllib来实现,正则表达式解析网页
    

    emmmmmmm

  • 相关阅读:
    搭建好lamp,部署owncloud。
    部署LAMP环境搭建一个网站论坛平台
    二进制编译安装httpd服务
    安装httpd服务并配置
    FTP的应用
    Linux配置IP,安装yum源
    RHEL-server-7.0-Linux-centos安装过程
    zabbix监控某一进程
    python获取windows系统的CPU信息。
    python相关cmdb系统
  • 原文地址:https://www.cnblogs.com/roccoshi/p/13027105.html
Copyright © 2011-2022 走看看