2020-10-22 Python爬虫第一章urllib库与requests库,第一节,深入学习urllib库

it2026-09-30  12

第一章  urllib库与requests库

第一节  urllib库

1.使用urlopen发起请求

爬虫概述。爬虫做的工作分为三步,第一步获取网页源代码,python提供了requests、urllib等库来获取网页源代码,第二步解析网页源代码,提取信息,在这里python提供了BeautifulSoup、pyquery、lxml等库来帮助我们提取自己想要的网页信息,第三步存储信息,在我们提取到信息后,有时候需要将信息保存在txt、json、csv等文件或者数据库中,这里数据库主要使用MySql,MongoDB数据库,若读者还会其他的数据库连接当然也可以。

1.1使用urlopen()函数发起请求,获取网页源代码。

使用urlopen(url)返回了一个Httpresponse对象,这个对象拥有许多方法,这里小编只列举一些常用的方法。

# from urllib.request import urlopen # urlopen(url,data=None,cafile = None,capath = None,cadefault = False,context = None,timeout = socket._GLOBAL_DEFAULT_TIMEOUT) # html = urlopen(url) # html.read() 将获取的网页的源代码以字节流的方式返回, # html.status 返回urlopen请求的状态码,以int型返回结果如200,301,404等 # html.getheaders() 返回响应头的信息,以列表形式返回结果 # html.close(),关闭urlopen请求,减少服务器的压力,假设你短时间内使用urlopen访问同一个网站上百次或上千次,若每次请求都没有关闭,则可能会报错,小编就曾经就遇到过一次,在使用close()后就再也没有报错

先来看看第一个案例,我们爬取的是http://pythonscraping.com/pages/page1.html网站,读者可以打开看看,为了便于观看,这里小编直接将结果解码了

from urllib.request import urlopen # 使用urlopen(url)函数发起请其余 html = urlopen('http://pythonscraping.com/pages/page1.html') #查看html返回类型 print(type(html)) content = html.read() print(type(content)) # 字节流解码 print(content.decode()) #------------------------------------------- ''' #以下是输出结果: #这里返回了urlopen的返回内省 <class 'http.client.HTTPResponse'> #输出结果为bytes类型 <class 'bytes'> <html> <head> <title>A Useful Page</title> </head> <body> <h1>An Interesting Title</h1> <div> Lorem ipsum dolor sit amet, consectetur adipisicing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum. </div> </body> </html> '''

我们再来看看urlopen()返回结果Httpresponse的其他方法,我们换一个网站看看,这里小编就省略了read方法了,毕竟这个网站内容太长。这次我们爬取http://pythonscraping.com/pages/warandpeace.html网站。

from urllib.request import urlopen #发起请求 html = urlopen('http://pythonscraping.com/pages/warandpeace.html') #返回请求状态码 print(type(html.status)) print(html.status) #返回请求头 print(html.getheaders()) #读者可自行执行 # print(html.read().decode()) html.close() #------------------------------------------- ''' #以下是输出结果 <class 'int'> 200 # 响应头信息 [('Date', 'Thu, 22 Oct 2020 13:50:58 GMT'), ('Server', 'Apache'), ('Last-Modified', 'Sat, 09 Jun 2018 19:15:59 GMT'), ('ETag', '"4121bd1-2dcb-56e3a58bcb54a"'), ('Accept-Ranges', 'bytes'), ('Content-Length', '11723'), ('Cache-Control', 'max-age=1209600'), ('Expires', 'Thu, 05 Nov 2020 13:50:58 GMT'), ('Connection', 'close'), ('Content-Type', 'text/html')] '''

1.2  urlopen函数的其他参数详解。     

在前面代码中我们介绍过urlopen()拥有许多参数,我们来说一下重要的三个参数url,timeout,data,url就不过多介绍了,这个是必不可少的参数,timeout设置请求的最长时间,当请求时间超过timeout,但这时还没有成功,则报错,data主要用于表单的提交(账号登陆就是表单提交中的一种),data参数只接受字节流值,后期将会介绍data的强大。我们来看看下面两个例子

from urllib.request import urlopen #第一个案例 #设置请求时间最长为2秒,如果你的网速慢可能这里也会报错,若报错则可将时间再调大试试 html = urlopen('http://pythonscraping.com/pages/warandpeace.html',timeout = 2) print(html.status) #设置请求时间最长为0.01秒 html_1 = urlopen('http://pythonscraping.com/pages/warandpeace.html',timeout = 0.01) #-------------------------------- ''' #以下是输出结果,第一个请求,请求成功,第二个失败 200 Traceback (most recent call last): ...(中间的错误小编省略了) socket.timeout: timed out ''' #第二个案例 from urllib.request import urlopen from urllib.parse import urlencode #由于网站特殊性,读者可设置任意字典 data = {'a':1,'b':'b'} #将urlencode(data)变为字节,后面会介绍urlencode data = bytes(urlencode(data),'utf-8') html = urlopen('http://httpbin.org/post',data = data) print(html.read().decode()) #----------------------------------- ''' #以下是输出结果,可以看出我们提供的字典,提交到了http://httpbin.org/post网上 { "args": {}, "data": "", "files": {}, "form": { "a": "1", "b": "b" }, "headers": { "Accept-Encoding": "identity", "Content-Length": "7", "Content-Type": "application/x-www-form-urlencoded", "Host": "httpbin.org", "User-Agent": "Python-urllib/3.6", "X-Amzn-Trace-Id": "Root=1-5f91927f-3981335d47fcb7cf45666b45" }, "json": null, "origin": "223.87.210.45", "url": "http://httpbin.org/post" } '''

2.使用urlopen+Request发起请求

2.1 使用urlopen+Request发起请求,修改请求头。    

通过上述方法我们能够获取一些普通的网站(如百度百科)的信息,但对于一些特殊网站,我们用上面方法,可能会获取不到信息。这些网站可能会通过识别你的请求头信息(上面输出的headers中,有一个是 "User-Agent": "Python-urllib/3.6"), 这种一看就知道是机器人请求(通过程序发起的请求)会被服务器拒之门外,所以我们要把我们的请求伪装成真人在浏览,其中最简单的就是修改请求头,下面我们来说使用Request来修改程序请求头。第一步,我们首先得获得请求头信息,具体步骤请参考http://cnblogs.com/Maple2cat/p/Python.html下的第一个headers的获取,打开自己想要爬取的网站,读者可将Accept,accept-language,accept-encode,connection,reffer,user_agent(最重要的一个必须要复制),host,cookie复制下来(有些时候cookie,reffer等参数可能不存在,则可将它忽略),也可以选择性复制,但user-agent必须要复制。至于这些参数啥作用读者可参考http://cnblogs.com/rxbook/p/9106446.html,只需查看request下对应的这8个参数就可以了,其他的看了也没啥用。下面我们看看Request修改请求头。可以看到我们程序的请求头被改变了,和上面程序输出的hearders就不一样,

from urllib.request import urlopen,Request #将你复制下来的东西,以字典形式存储,像小编这样,这里小编只复制了7个 headers = {'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Connection': 'keep-alive', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36', 'Host': 'www.baidu.com', 'Referer': 'https://mp.csdn.net/console/editor/html/109230346'} #http://httpbin.org/get返回你的请求头信息 r = Request('http://httpbin.org/get',headers = headers) html = urlopen(r) print(html.read().decode()) #------------------------------------- ''' # 以下是输出结果 { "args": {}, "headers": { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "zh-CN,zh;q=0.9", "Host": "www.baidu.com", "Referer": "https://mp.csdn.net/console/editor/html/109230346", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36", "X-Amzn-Trace-Id": "Root=1-5f9199d4-6c4f46ab2b744cc03d95d8bb" }, "origin": "223.87.210.45", "url": "http://www.baidu.com/get" } '''

2.2我们使用cookie+Request实现直接登陆网站

读者先使用浏览器打开这个链接http://login2.scrape.cuiqingcai.com/login,用户名:admin,密码:admin,我们以这网站为例

在登陆进去之后,和获取请求头headers方法一样,只需复制我上面所说的参数复制下来就可以了,其他的参数(如:authority)不用复制,当然我们所说的你也可以选择性复制,但cookie,user_agent必须复制。注意若直接复制小编的代码运行,是运行不了的,因为这个网站cookie是变化的,所以读者必须自行登录,将cookie更改后才能运行,并且注意代码中小编给的注意事项,不然也会报错

from urllib.request import Request,urlopen headers = {'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', #注意事项 #读者可试试将accept-encoding,放进去,若放进去我们可能得不到我们想要的结果,他甚至报错 #对于这一项内容,小编建议能不放就不放,说多了都是泪 # 'accept-encoding': 'gzip, deflate, br', 'accept-language': 'zh-CN,zh;q=0.9', 'cache-control': 'max-age=0', 'cookie': 'UM_distinctid=1750cb8533a6ed-037f7cb9214222-5040231b-144000-1750cb8533b813; Hm_lvt_3ef185224776ec2561c9f7066ead4f24=1602236208,1602498074,1602498302,1602499124; sessionid=2pped7lo3cg1se9nwx2isc3se52g4khh', 'referer': 'https://login2.scrape.cuiqingcai.com/login', 'user-agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36'} #注意事项 #注意这里cookie是加密文本,所以这里只能使用https协议,若使用http协议则会弹出308错误, # 若读者复制连接下来为http,也要改为https r = Request('https://login2.scrape.cuiqingcai.com',headers = headers) html = urlopen(r) #我们就获取登陆后的信息了 print(html.read().decode()) #--------------------------------------- ''' #下面是部分结果,由于结果太长,我只发了中间部分结果 <img data-v-7f856186="" src="https://p0.meituan.net/movie/ce4da3e03e655b5b88ed31b5cd7896cf62472.jpg@464w_644h_1e_1c" class="cover"> </a> </div> <div data-v-7f856186="" class="p-h el-col el-col-24 el-col-xs-9 el-col-sm-13 el-col-md-16"> <a data-v-7f856186="" href="/detail/1" class="name"> <h2 data-v-7f856186="" class="m-b-sm">霸王别姬 - Farewell My Concubine</h2> </a> '''

 2.3 详解Request其他参数

from urllib.request import Request,urlopen r = Request(url,data = None,headers = None,origin_req_host = None,unverifiable = False,method = 'GET') #请求 可将r视为一个链接传入 html = urlopen(r) #在众多参数中我们需要掌握前三个和最后一个 # url 表示请求连接(你需要爬取的网站) # data 你爬取网站所需要提交的信息,比如你的账号,密码,验证码等,以字节流方式传入 # headers 设置自己程序的请求头,以字典方式传入 # origin_req_host 设置请求方的IP地址,又是我们需要设置其他的ip就可以通过这个方法 # method 表示程序的请求方式,通常有GET,POST,DELETE,PUT. 具体差距见下 #后面两个只需了解就可以了,我们重点掌握前两个 # GET方法可以理解为告诉服务器。请按照这个地址给我信息,一般我们在浏览器中输入连接再按回车方式,都是get请求 # POST 可以理解为告诉服务器,帮我把这个数据存储到你的数据库里,比如我们在登陆网站输出密码账号时,用的都是post方法 # PUT可以理解为告诉服务器,帮我把我的信息个按这个更新了, # DELETE主要用于服务器删除某个对象

 通过上面两节的内容,我们知道如下两个重要内容,第一使用urlopen向一些普通网站(百度百科等)发起请求,并输出内容,第二使用urlopen+Request发起请求,重点是学会修改自己的请求头,若已有已经登陆网站(这方法可能只对部分网站有用)的cookie我们如何去实现登陆。

3.build_opener发起请求

3.1 了解build_opener。

有时候我们知道网站的账号以及密码,那我们如何实现登录呢,下面我们将学习urllib下的重要build_opener函数(build_opener函数返回Opener类,顾后面小编说build_opener函数,build_opener函数都特指Opener类),这个类的作用比我们urlopen+Request还要强大。我们之前使用的urlopen方法,它就是通过build_opener函数来实现的,urlopen()的返回类型和build_opener().open()返回的类型一致,这就说明,urlopen的其它方法, build_opener().open()也能使用

我们还是来看是否属实

from urllib.request import build_opener #实例化build_opener(),因为这个函数返回去一个类Opener opener = build_opener() #调用opener的方法open html = opener.open('http://pythonscraping.com/pages/page1.html') #查看来类型 print(type(html)) print(html.read().decode()) #-------------------------------------- ''' #以下是输出结果 <class 'http.client.HTTPResponse'> <html> <head> <title>A Useful Page</title> ...(为了缩短篇幅,小编省略了) </body> </html>

这也就说明了,我们可以使用opener.open()时,我们可以使用Httpresponse的方法html.status,html.getheaders(),html.close(),html.url(表示返回请求的链接,小编上面忘记说了,urlopen也有html.url这个属性)

3.2 build_opener修改请求头。   

我们再来说说build_opener函数的其它功能,第一个功能修改请求头我们还是用http://httpbin.org/get网站,读者可以试试若不修改请求头,会是啥样的。

from urllib.request import build_opener opener = build_opener() #注意opener的修改headers,是使用赋值方式,而且headers的传入方式是以列表形式传入,而不是字典 headers = [('Accept', 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9'), ('Accept-Encoding', 'gzip, deflate, br'), ('Accept-Language', 'zh-CN,zh;q=0.9'), ('Connection', 'keep-alive'), ('User-Agent', 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36'), ('Host', 'www.baidu.com'), ('Referer', 'https://mp.csdn.net/console/editor/html/109230346')] opener.addheaders = headers html = opener.open('http://httpbin.org/get') print(html.read().decode()) #---------------------------------- ''' #以下时输出结果 { "args": {}, "headers": { "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9", "Accept-Encoding": "gzip, deflate, br", "Accept-Language": "zh-CN,zh;q=0.9", "Host": "httpbin.org", "Referer": "https://mp.csdn.net/console/editor/html/109230346", "User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36", "X-Amzn-Trace-Id": "Root=1-5f92453e-140170521667b41a664ab8de" }, "origin": "223.87.208.38", "url": "http://httpbin.org/get" } '''

可以看到我们的请求头成功被修改,而且是不依靠Request类,当然opener也能使用Request修改请求头,方式见下面

from urllib.request import build_opener,Request #我们是将headers放入Request中,故我们需要将headers设置为字典形式,而不是列表 headers = {'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9', 'Accept-Encoding': 'gzip, deflate, br', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Connection': 'keep-alive', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.79 Safari/537.36', 'Host': 'www.baidu.com', 'Referer': 'https://mp.csdn.net/console/editor/html/109230346'} r = Request('http://httpbin.org/get',headers = headers) opener = build_opener() html = opener.open(r) print(html.read().decode()) #输出结果与上一段代码一致,小编就不列举了

3.2  实现账号登陆。   

 现在我们来说说,如何使用build_opener来向网站发送账号密码等,我们还是以http://login2.scrape.cuiqingcai.com/login为例

from urllib.request import build_opener,HTTPCookieProcessor from http.cookiejar import CookieJar from urllib.parse import urlencode #设置一个存储cookies的容器cookie,前面三句几乎是固定的 cookie = CookieJar() handler = HTTPCookieProcessor(cookie) opener = build_opener(handler) #将账号密码放入字典,再将其变为字节流 data = {'username':'admin','password':'admin'} data = bytes(urlencode(data),'utf-8') #这个链接是你登录界面url,传入参数。 html = opener.open('https://login2.scrape.cuiqingcai.com/login',data = data) #看看cookie中存的到底是什们东西 print(cookie) #在第一个请求成功后,opener就会将cookies存入cookie中,之后的链接请求若需要cookies,则opener直接调用cookie中的cookies #这个链接是你登陆成功之后需要访问的链接 html_1 = opener.open('https://login2.scrape.cuiqingcai.com') print(html_1.read().decode()) #------------------------------- ''' #以下是输出部分结果 <CookieJar[<Cookie sessionid=py4kod3c7dysbqqh147i481j037yhme2 for login2.scrape.cuiqingcai.com/>]> ... <img data-v-7f856186="" src="https://p0.meituan.net/movie/ce4da3e03e655b5b88ed31b5cd7896cf62472.jpg@464w_644h_1e_1c" class="cover"> </a> </div> <div data-v-7f856186="" class="p-h el-col el-col-24 el-col-xs-9 el-col-sm-13 el-col-md-16"> <a data-v-7f856186="" href="/detail/1" class="name"> <h2 data-v-7f856186="" class="m-b-sm">霸王别姬 - Farewell My Concubine</h2> ... '''

当然我们也能使用opener+Request模拟登陆这里我们就换一个网站为例

http://pythonscraping.com/pages/cookies/login.html用户名任意字段,密码为password

from urllib.request import build_opener,HTTPCookieProcessor,Request from http.cookiejar import CookieJar from urllib.parse import urlencode #设置一个存储cookies的容器cookie,前面三句几乎是固定的 cookie = CookieJar() handler = HTTPCookieProcessor(cookie) opener = build_opener(handler) #将账号密码放入字典,再将其变为字节流 data = {'username':'a','password':'password'} data = bytes(urlencode(data),'utf-8') #这个网站不是重定向网站,故需要在第二个界面发起请求 r = Request('https://pythonscraping.com/pages/cookies/welcome.php',data = data) html = opener.open(r) #访问登陆成功后的页面 html_1 = opener.open('https://pythonscraping.com/pages/cookies/profile.php ') print(html_1.read().decode()) #若需要发起查看在成功登录后第一个请求下网页,需要重新发起请求 html_2 = opener.open('https://pythonscraping.com/pages/cookies/welcome.php') print(html_2.read().decode()) #---------------------------------- ''' #以下是输出结果 Hey a! Looks like you're still logged into the site! <h2>Welcome to the Website!</h2> You have logged in successfully! <br><a href="profile.php">Check out your profile!</a> '''

我们除了用上述的方法外,若我们知道登陆后的cookies,我们也可以将将cookies放入headers,使用build_opener的两种添加请求头的方式,就能够顺利登陆。下面附上了open全部参数,这些参数原理和urlopen一致,若有不理解则可翻看前面解说

opener = build_opener() html = opener.open(url,data = None,timeout = socket._GLOBAL_DEFAULT_TIMEOUT)

至于使用urllib做Ip代理,对于新手来说只要不去频繁爬取大型的网站,本地IP就可以用了。而且小编现在手上也没有可用的IP,就算有多余的IP号,我即便附上IP后,可能过不久这个IP就会被封了。大家也试不了。后期我们讲IP代理池时,再仔细深入。

3.4 urllib的其它重要参数

urllib是一个很强大的库他下面拥有很多函数。帮我们解决链接的各种处理。见下

from urllib.parse import urlparse,urlunparse,urlsplit,urlunsplit,urlencode,urljoin,parse_qs,parse_qsl,quote,unquote #有十个处理链接的函数,不过这些中大部分可以使用正则表达式帮我们处理,我们只需了解就可以 #重点掌握urlencode,quote就可以了 # 下面我们来说说quotem,urlencode # urlencode(a),将a对象处理后返回为一个字符串对象。当然我们对a是有要求它必须是一个字典或者[(a1,b1],(a2,b2)]这样格式的才能行,就下面的案例 #原官方对a的要求是这样的Encode a dict or sequence of two-element tuples 我觉得和我解释的差不多 a = [('a',2)] a = [(1,2)] print(urlencode(a)) print(type(urlencode(a))) #--------------------------- '''输出结果 a = 2 1 = 2 <class str> ''' #quote主要处理链接中有中文的url #主要搭配string使用 import string #将s看做一颗经过特殊处理的url s = quote(url,safe = string.printable ) html = urlopen(s) #或者 r = Request(s) html = urlopen(r)

到这为止,我们就差不多讲完了第一个库urllib的学习,这一个库学习分为三个阶段。

第一阶段使用urlopen发起基本的请求。

第二阶段使用urlopen+Request修改请求头,以及使用已有的cookies,实现简单的登陆。

第三阶段使用build_opener,修改请求头,以及使用已有的cookies,实现简单的登陆,和使用账号进行登陆。

下一节我们将说更为强的库requests。

我们可以理解为它是urllib的一个进阶版,它的语法比urllib的简单的多了。

#小编是一个理科生,文笔不好,大家能理解我的意思就行了

#转载则请标明文章出处,谢谢

#文中若有任何错误,欢迎大家积极指出,小编洗耳恭听

最新回复(0)