网络爬虫是从公开网页采集数据的自动化技术。本文覆盖从单页请求到分布式爬虫、从静态页面到浏览器自动化的完整工具箱,并强调合规边界。
目录
- 爬虫的合规边界
- HTTP 请求:requests 与 httpx
- 解析 HTML:BeautifulSoup 与 lxml
- XPath 与 CSS 选择器
- JSON API 与异步抓取
- Scrapy 框架
- Playwright 浏览器自动化
- 反爬应对与限速
- 数据清洗与存储
- 速查表与最佳实践
1. 爬虫的合规边界
在动手之前,先确认边界:
| 事项 | 说明 |
|---|---|
robots.txt | 网站声明的抓取规则(robots.txt 里 Disallow 的路径别抓) |
| 服务条款 | 网站 ToS 是否禁止抓取 |
| 数据用途 | 个人学习/研究 vs 商业再分发 |
| 频率控制 | 别压垮目标服务器 |
| 登录数据 | 不要采集非公开数据 |
# 尊重 robots.txt
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
can = rp.can_fetch("my-bot/1.0", "https://example.com/private/")
print("允许抓取?", can)
提醒:技术可行 ≠ 合法合规。批量抓取应优先考虑网站是否提供官方 API。
2. HTTP 请求:requests 与 httpx
2.1 requests 基础
import requests
session = requests.Session() # 复用连接 + 保持 cookie
session.headers.update({"User-Agent": "my-scraper/1.0"})
resp = session.get("https://example.com/list?page=1", timeout=10)
print(resp.status_code, resp.url)
html = resp.text # 默认按 header 解码
print(resp.encoding) # 可能是 ISO-8859-1,需要处理中文
resp.encoding = "utf-8" # 修正编码
2.2 httpx(同步/异步双模式)
import httpx
# 同步
with httpx.Client(headers={"User-Agent": "x"}, timeout=10) as client:
r = client.get("https://example.com")
# 异步(高并发)
import asyncio
async def main():
async with httpx.AsyncClient(timeout=10) as client:
tasks = [client.get(f"https://example.com/item/{i}") for i in range(50)]
results = await asyncio.gather(*tasks)
for r in results:
if r.status_code == 200:
process(r.text)
asyncio.run(main())
2.3 处理分页与循环
def crawl_pages(base_url, max_pages=20):
session = requests.Session()
for page in range(1, max_pages + 1):
resp = session.get(f"{base_url}?page={page}", timeout=10)
if resp.status_code != 200:
break
yield parse_page(resp.text)
time.sleep(0.5) # 限速,别打太狠
3. 解析 HTML:BeautifulSoup 与 lxml
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "lxml") # lxml 快;html.parser 更稳
# 提取标题
print(soup.title.text)
# 找单个元素
title = soup.find("h1")
# 按 class
items = soup.find_all("div", class_="product-item")
# 按属性
links = soup.find_all("a", href=True)
# CSS 选择器(更强大)
cards = soup.select("div.card > a.title")
for item in items:
name = item.select_one("h2").text.strip()
price = item.select_one(".price").text.strip()
link = item.select_one("a")["href"]
print(name, price, link)
lxml 直接解析(更快):
from lxml import html as lxml_html
tree = lxml_html.fromstring(html)
# XPath
prices = tree.xpath("//div[@class='price']/text()")
4. XPath 与 CSS 选择器
| 目标 | CSS | XPath |
|---|---|---|
| id 选择 | #price | //*[@id='price'] |
| class | .card | //*[contains(@class,'card')] |
| 子元素 | div > a | //div/a |
| 属性 | a[href^='http'] | //a[starts-with(@href,'http')] |
| 文本 | — | //h1/text() |
| 第 n 个 | li:nth-child(2) | //li[2] |
# XPath 实战:提取所有商品链接与标题
links = tree.xpath("//div[@class='product']//a[@class='title']")
for a in links:
text = a.text_content().strip()
href = a.get("href")
print(text, href)
调试选择器:浏览器 F12 → 复制 selector / XPath,再在 Python 里验证。
5. JSON API 与异步抓取
很多「网页」背后是 JSON API,直接调 API 更高效:
import httpx, asyncio, json
async def crawl_api():
async with httpx.AsyncClient() as client:
# 分页 JSON API
all_items = []
for page in range(1, 5):
r = await client.get(
"https://example.com/api/items",
params={"page": page, "limit": 100},
)
data = r.json()
all_items.extend(data["items"])
if page >= data["total_pages"]:
break
# 并发抓详情
async def fetch_detail(item_id):
r = await client.get(f"https://example.com/api/items/{item_id}")
return r.json()
details = await asyncio.gather(
*(fetch_detail(i["id"]) for i in all_items[:20])
)
return details
data = asyncio.run(crawl_api())
print(json.dumps(data, ensure_ascii=False, indent=2)[:500])
优点:请求少、数据干净、不受 HTML 改版影响。优先找 API。
6. Scrapy 框架
大规模爬虫用 Scrapy:自带调度、去重、管道、并发、中间件。
# items.py
import scrapy
class ProductItem(scrapy.Item):
name = scrapy.Field()
price = scrapy.Field()
url = scrapy.Field()
# spiders/products.py
import scrapy
class ProductSpider(scrapy.Spider):
name = "products"
start_urls = ["https://example.com/products"]
def parse(self, response):
for card in response.css("div.product"):
item = ProductItem()
item["name"] = card.css("h2::text").get()
item["price"] = card.css(".price::text").get()
item["url"] = card.css("a::attr(href)").get()
yield item
# 自动翻页
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, self.parse)
# 运行
scrapy crawl products -O products.json
Scrapy 核心优势:
| 特性 | 说明 |
|---|---|
| 去重 | 内置 URL 去重 |
| 并发 | 自动调度异步请求 |
| 管道 | 数据清洗/存储流水线 |
| 中间件 | 代理、限速、UA 轮换 |
| 断点续爬 | JOBDIR |
7. Playwright 浏览器自动化
需要执行 JavaScript 的页面(SPA、动态渲染)用 Playwright:
pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/dashboard", wait_until="networkidle")
page.wait_for_selector(".data-table") # 等元素渲染
# 交互:点击、输入、滚动
page.click("#load-more")
page.fill("#search", "python")
page.wait_for_timeout(1000)
# 提取
rows = page.query_selector_all("tr")
for row in rows[:5]:
print(row.inner_text())
# 截图 / PDF(便于调试)
page.screenshot(path="debug.png")
browser.close()
异步版 + 并发浏览器:
from playwright.async_api import async_playwright
import asyncio
async def snap(url):
async with async_playwright() as p:
b = await p.chromium.launch()
pg = await b.new_page()
await pg.goto(url)
title = await pg.title()
await b.close()
return title
async def main():
titles = await asyncio.gather(*[snap(u) for u in URLS])
print(titles)
8. 反爬应对与限速
8.1 常见反爬手段与对策
| 反爬手段 | 对策 |
|---|---|
| UA 校验 | 设置真实浏览器 UA |
| 频率限制 | 限速 + 指数退避重试 |
| IP 封禁 | 代理池轮换 |
| JS 动态渲染 | Playwright |
| 验证码 | 人工兜底 / 服务降级 |
| 数据加密 | 逆向或转 API |
# 指数退避
def fetch_with_retry(url, max_retries=3):
for i in range(max_retries):
try:
return requests.get(url, timeout=10)
except requests.RequestException as e:
wait = 2 ** i
print(f"重试 {i+1},等待 {wait}s: {e}")
time.sleep(wait)
raise RuntimeError(f"最终失败: {url}")
8.2 限速设计(礼貌爬虫)
import time, random
def polite_fetch(url, min_delay=1.0, max_delay=3.0):
time.sleep(random.uniform(min_delay, max_delay)) # 随机延时更像人
return requests.get(url, timeout=10)
8.3 代理池
proxies = [
"http://proxy1:8080",
"http://proxy2:8080",
]
def get_proxy():
return {"http": random.choice(proxies), "https": random.choice(proxies)}
态度:爬虫要保持「礼貌」——限速、带 UA、尊重 robots、失败退避。封禁往往源于抓太狠而非抓本身。
9. 数据清洗与存储
9.1 清洗管道
def clean_price(raw: str) -> float:
"""'¥1,299.00' → 1299.00"""
import re
s = re.sub(r"[^\d.]", "", raw)
return float(s) if s else 0.0
def clean_text(raw: str) -> str:
return " ".join(raw.split()) # 合并多余空白
# 结构化
records = [
{"name": clean_text(name), "price": clean_price(price)}
for name, price in raw_items
]
9.2 存储选项
| 数据量 | 方案 |
|---|---|
| 小 | JSONL / CSV |
| 中 | SQLite |
| 大 | PostgreSQL / Parquet |
| 需要变更追踪 | 追加 JSONL + 定期去重 |
import json, sqlite3
# JSONL 追加
with open("items.jsonl", "a", encoding="utf-8") as f:
f.write(json.dumps(record, ensure_ascii=False) + "\n")
# SQLite
conn = sqlite3.connect("items.db")
conn.execute("CREATE TABLE IF NOT EXISTS items (name TEXT, price REAL)")
conn.execute("INSERT INTO items VALUES (?, ?)", (record["name"], record["price"]))
conn.commit()
10. 速查表与最佳实践
| 场景 | 工具 |
|---|---|
| 简单静态页 | requests + BeautifulSoup |
| 大规模并发 | httpx 异步 / Scrapy |
| JSON API | requests / httpx 直接调 |
| JS 动态页 | Playwright |
| 快速实验 | requests + lxml |
| 数据量大 | Scrapy + PostgreSQL |
最佳实践清单:
- 先看 robots.txt + 找官方 API。
- 用
Session/Client复用连接。 - 设置超时、限速、退避重试。
- 统一数据 schema,清洗函数复用。
- 抓取结果落盘(JSONL),断点续抓。
- 出错要可恢复,别一崩全丢。
- 遵守目标网站 ToS,个人使用为主。
一句话记忆:爬虫四件套——requests 发请求、BeautifulSoup/lxml 解析、Scrapy 管规模、Playwright 过动态;永远先问「有没有 API、让不让抓」。
延伸阅读
- Python 网络编程:socket、HTTP 与服务端 —— HTTP 客户端底层
- Python 文件 IO 与数据序列化 —— 数据存储格式
- Python 异步高级 —— 异步爬虫并发
- [[network]] —— HTTP 协议与代理原理
- [[security]] —— 爬虫安全与合规
爬虫是数据工程的起点,也是与 Web 生态互动的窗口。会抓只是第一步,会「干净地抓、合规地抓、可持续地抓」才是工程能力。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。