# 靜態網頁的資料爬蟲策略


* 了解靜態網頁的資料爬蟲策略
* 認識適用於靜態網頁爬蟲的相關套件工具：Request
* 認識適用於靜態網頁爬蟲的相關套件工具：BeatifulSoup

## 作業目標

利用 Request + BeatifulSoup 爬取下列兩個網站內容並解析：

1. Dcard 網址： https://www.dcard.tw/f
2. 知乎： https://www.zhihu.com/explore

並且回答下面問題：

1. Request 取回之後該怎麼取出資料，資料型態是什麼？
2. 為什麼要使用 BeatifulSoup 處理？處理後的型態是什麼？
3. 觀察一下知乎回來的資料好像有點怪怪的，該怎麼解決？

In [2]:
import requests
from bs4 import BeautifulSoup


In [5]:
url = 'https://www.dcard.tw/f'
r = requests.get(url)
r.encoding = 'utf-8'
print(r.text[0:3000])

<!DOCTYPE html><html lang="zh-Hant-TW"><head prefix="og: http://ogp.me/ns#" itemscope="" itemType="https://schema.org/WebSite"><title data-react-helmet="true">Dcard</title><meta data-react-helmet="true" property="og:image" content="https://www.dcard.tw/build/landing-c9e7b8fb.png"/><meta data-react-helmet="true" property="og:image:secure_url" content="https://www.dcard.tw/build/landing-c9e7b8fb.png"/><meta data-react-helmet="true" charSet="utf-8"/><meta data-react-helmet="true" http-equiv="X-UA-Compatible" content="IE=edge"/><meta data-react-helmet="true" name="application-name" content="Dcard"/><meta data-react-helmet="true" name="apple-itunes-app" content="app-id=951353454"/><meta data-react-helmet="true" name="theme-color" content="#006aa6"/><meta data-react-helmet="true" name="mobile-web-app-capable" content="yes"/><meta data-react-helmet="true" name="apple-mobile-web-app-capable" content="yes"/><meta data-react-helmet="true" property="fb:app_id" content="211628828926493"/><meta dat

In [6]:
print('Request 取回之後該怎麼取出資料，資料型態是什麼？ => ', type(r.text))

Request 取回之後該怎麼取出資料，資料型態是什麼？ =>  <class 'str'>


In [10]:
soup = BeautifulSoup(r.text, "html5lib")

print(soup)

<!DOCTYPE html>
<html lang="zh-Hant-TW"><head itemscope="" itemtype="https://schema.org/WebSite" prefix="og: http://ogp.me/ns#"><title data-react-helmet="true">Dcard</title><meta content="https://www.dcard.tw/build/landing-c9e7b8fb.png" data-react-helmet="true" property="og:image"/><meta content="https://www.dcard.tw/build/landing-c9e7b8fb.png" data-react-helmet="true" property="og:image:secure_url"/><meta charset="utf-8" data-react-helmet="true"/><meta content="IE=edge" data-react-helmet="true" http-equiv="X-UA-Compatible"/><meta content="Dcard" data-react-helmet="true" name="application-name"/><meta content="app-id=951353454" data-react-helmet="true" name="apple-itunes-app"/><meta content="#006aa6" data-react-helmet="true" name="theme-color"/><meta content="yes" data-react-helmet="true" name="mobile-web-app-capable"/><meta content="yes" data-react-helmet="true" name="apple-mobile-web-app-capable"/><meta content="211628828926493" data-react-helmet="true" property="fb:app_id"/><meta co

In [11]:
print('為什麼要使用 BeatifulSoup 處理？處理後的型態是什麼？ => ', type(soup))

為什麼要使用 BeatifulSoup 處理？處理後的型態是什麼？ =>  <class 'bs4.BeautifulSoup'>


In [12]:
url = 'https://www.zhihu.com/explore'
r = requests.get(url)
r.encoding = 'utf-8'
print(r.text[0:600])

<html>
<head><title>400 Bad Request</title></head>
<body bgcolor="white">
<center><h1>400 Bad Request</h1></center>
<hr><center>openresty</center>
</body>
</html>



## 觀察一下知乎回來的資料好像有點怪怪的，該怎麼解決？

In [13]:
url = 'https://www.zhihu.com/explore'

headers = {
    'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3',
    'accept-encoding': 'gzip, deflate, br',
    'accept-language': 'zh-TW,zh;q=0.9,en-US;q=0.8,en;q=0.7,zh-CN;q=0.6,ja;q=0.5',
    'cache-control': 'max-age=0',
    'cookie': '_zap=a5634a47-a806-4828-8ab4-2cbc03734bda; d_c0="AFBoGjc1fg6PTus2_H76YgMM3xvztVHRnCs=|1541767835"; __gads=ID=18e1196642fc5994:T=1544975081:S=ALNI_MZ2lQjiLHPlLxjVqGGH6o-EiL2luQ; z_c0="2|1:0|10:1551204934|4:z_c0|92:Mi4xM1FkMERnQUFBQUFBVUdnYU56Vi1EaVlBQUFCZ0FsVk5SdEJpWFFDUXZ6REVHVWw0QWtXblh6WGt4T183ekdxNFhn|fa542d11758c34207cdd6a1edf85de768ac02624c8cb53aa687017af49accd7b"; tst=r; q_c1=aed50b9b158344d6ac78a230c8970d83|1560948516000|1543510922000; __utmv=51854390.100--|2=registration_date=20190226=1^3=entry_date=20181130=1; _xsrf=13cd498d-6139-4c50-99a6-df5c749dd64c; tgw_l7_route=4860b599c6644634a0abcd4d10d37251; __utma=51854390.567487929.1560948518.1560950333.1562469680.3; __utmb=51854390.0.10.1562469680; __utmc=51854390; __utmz=51854390.1562469680.3.3.utmcsr=localhost:8888|utmccn=(referral)|utmcmd=referral|utmcct=/notebooks/day2-example.ipynb',
    'referer': 'http://localhost:8888/notebooks/day2-example.ipynb',
    'upgrade-insecure-requests': '1',
    'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_6) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/75.0.3770.100 Safari/537.36'
}
r = requests.get(url, headers=headers)

r.encoding = 'utf-8'
print(r.text[0:600])

<!DOCTYPE html>
<html lang="zh-CN" dropEffect="none" class="no-js ">
<head>
<meta charset="utf-8" />

<meta http-equiv="X-UA-Compatible" content="IE=edge,chrome=1" />
<meta name="renderer" content="webkit" />
<meta http-equiv="X-ZA-Response-Id" content="a86da3a4ed72457d">

<meta http-equiv="X-ZA-Experiment" content="default:None,ge3:ge3_9,ge2:ge2_1,SE_I:c,nwebQAGrowth:experiment,recom_answer_rec:answer_gbdt_rank,is_office:false,nweb_growth_people:default,app_store_rate_dialog:close,live_store:ls_a3_b2_c2_f2,nweb_search:nweb_search_heifetz,new_live_feed_mediacard:new,hybrid_zhmore_video:yes,ad_
