目標:登錄大話西游論壇入口網站http://dhxy.netease.com/forum-39-1.html 自定義頁數爬取
這個版塊
版主回復
工具:scrapy框架
思路:獲取start_urls 網頁內容,獲取帖子鏈接,存入數組,將帖子鏈接一個一個傳遞給parse_content 函數解析內容,parse_content生成字典并且獲取下一頁鏈接,回調parse函數,重新進行下一頁的爬取
代碼實現:
創建項目:scrapy startporject dhxy_luntan
根據提示創建 spider
spider.py
# -*- coding: utf-8 -*-
import scrapy
import re
import sys
from scrapy.http import Request
from dhxy_luntan.items import DhxyLuntanItem
reload(sys)
sys.setdefaultencoding('utf-8')
class DhxyLuntanSpiderSpider(scrapy.Spider):
name = "dhxy_luntan_spider"
#allowed_domains = ["http://dhxy.netease.com/forum-39-1.html"]
start_urls = (
'http://dhxy.netease.com/forum-39-1.html',
)
def parse(self, response):
urls = []
sel = scrapy.Selector(response)
content = sel.xpath('//*[@id="threadlist"]/div[2]/form/table/tbody[starts-with(@id,"normalthread_")]/tr/th').extract()
print len(content)
str_ = u'版主回復'
qiandao_ = u'簽到'
for i in content:
if str_ in i:
hrefs = re.findall('</em> <a href="(.*?)" .*?onclick.*?</a>', i, re.S)
if hrefs:
href = hrefs[0]
full_ + href
urls.append(full_href)
for i in urls:
yield Request(i, callback=self.parse_content)
def parse_content(self, response):
item = DhxyLuntanItem()
sel = scrapy.Selector(response)
title = sel.xpath('//*[@id="thread_subject"]/text()').extract()[0]
item['url'] = u'\n['+title + u']' + response.url
data = sel.xpath('//tr/td[starts-with(@id, "postmessage")]')
content = data.xpath('string(.)').extract() #string(.) 當前層的所有內容作為一個字符串輸出
if content:
item['content'] = content
yield item
for i in range(2,5):
next_page = 'http://dhxy.netease.com/forum-39-' + str(i) + '.html'
yield scrapy.http.Request(next_page, callback = self.parse)
items.py
Paste_Image.png
setting.py
因為需要輸出csv格式,在setting文件中設置
Paste_Image.png
pipeline.py
不用保存到數據庫pipeline無需設置
運行代碼:scrapy crawl dhxy_luntan_spider
生成csv文件:
Paste_Image.png
使用excel的導出功能,更改文件類型:
Paste_Image.png
特殊字符可在編輯器中替換,?分析為 (html 里是空格占位符,普通的空格在 html 里如果連續的多個可能被認為只有一個,而這個東西你寫幾個就能占幾個空格位)