zoukankan html css js c++ java

BeautifulSoup 抓取网站url

  1 # -*- coding:utf-8 -*-
  2 import urlparse
  3 import urllib2
  4 from bs4 import BeautifulSoup
  5 
  6 url = "http://www.baidu.com"
  7 
  8 urls = [url] # stack of urls to scrape
  9 visited = [url] # historic record of urls
 10 
  1 # -*- coding:utf-8 -*-
  2 import urlparse
  3 import urllib2
  4 from bs4 import BeautifulSoup
  5 
  6 url = "http://www.baidu.com"
  7 
  8 urls = [url] # stack of urls to scrape
  9 visited = [url] # historic record of urls
 10 
 11 while len(urls) > 0:
 12     try:
 13         htmltext = urllib2.urlopen(urls[0]).read()
 14     except:
 15         print urls[0]
 16     soup = BeautifulSoup(htmltext,"html")
 17 
 18     urls.pop(0)
 19 
 20     for tag in soup.findAll("a", href=True):
 21         tag["href"] = urlparse.urljoin(url, tag["href"])
 22         if url in tag["href"] and tag["href"] not in visited:
 23             urls.append(tag["href"])
 24             visited.append(tag["href"])
 25 
 26     print len(urls)

查看全文

相关阅读:
采购标准流程及底层分析
 ORACLE FORM ZA 常用子程序
 在R12中实现多OU编程
 FORM未找到数据的原因
 在Oracle的FORM中高亮显示鼠标点击或光标所在的行
 MPICH运行程序时出错之解决方法
 两个基于C++的MPI编辑例子
 面向对象PHP面向对象的特性
 PHP 数组遍历 foreach 语法结构
 php BC高精确度函数库

原文地址：https://www.cnblogs.com/cuzz/p/BeautifulSoup.html