目前正在學Python爬蟲,正在讀崔慶才的《Python3網(wǎng)絡爬蟲開發(fā)實戰(zhàn)》,之前學習正則表達式,但是由于太難,最后放棄了(學渣的眼淚。。。。),在這本書上的抓取貓眼電影排行上,后來自學了pyquery,發(fā)現(xiàn)用pyquery可以解決這個問題,目前自己試著寫了代碼
附上代碼:
import requests
from pyquery import PyQuery as pq
import time
def get_one_page(url):
headers = {
'User-Agent':'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.79 Safari/537.36'
}
html = requests.get(url=url,headers=headers)
return html.text
def parse_one_page(html):
doc = pq(html)
items = doc('dd').items()
for item in items:
item1 = item.find('.board-item-main .board-item-content .movie-item-info')#空格表示嵌套
item2 = item.find('.board-index')
print('名次:' + item2.text())
name = item1.find('.name').text()
star = item1.find('.star').text()
time = item1.find('.releasetime').text()
score = item1.siblings('.movie-item-number .score .integer').text() + item1.siblings('.movie-item-number .score .fraction').text()
print('電影名:' + name + '\n' +
star + '\n' + time + '\n' + '評分:'+score +'\n')
def main(offset):
url = 'http://maoyan.com/board/4?offset=' + str(offset) #設置偏移量
html = get_one_page(url)
parse_one_page(html)
if __name__ == '__main__':
for i in range(10):
main(offset = i * 10)
time.sleep(1)#由于現(xiàn)在貓眼多了反爬蟲,如果速度過快則無響應,所以要添加延時等待。
代碼在:
https://github.com/liuweixu/Python-crawler/tree/master/PyQuery
歡迎來star