Python 网络爬虫与自动化:从 requests 到 Playwright

Python 网络爬虫与自动化完整指南:requests/httpx 请求、BeautifulSoup/lxml 解析、Scrapy 框架、Playwright 浏览器自动化、反爬应对(限速/代理/指纹)、robots 合规与数据清洗工程。

网络爬虫是从公开网页采集数据的自动化技术。本文覆盖从单页请求到分布式爬虫、从静态页面到浏览器自动化的完整工具箱,并强调合规边界。


目录

  1. 爬虫的合规边界
  2. HTTP 请求:requests 与 httpx
  3. 解析 HTML:BeautifulSoup 与 lxml
  4. XPath 与 CSS 选择器
  5. JSON API 与异步抓取
  6. Scrapy 框架
  7. Playwright 浏览器自动化
  8. 反爬应对与限速
  9. 数据清洗与存储
  10. 速查表与最佳实践

1. 爬虫的合规边界

在动手之前,先确认边界:

事项说明
robots.txt网站声明的抓取规则(robots.txt 里 Disallow 的路径别抓)
服务条款网站 ToS 是否禁止抓取
数据用途个人学习/研究 vs 商业再分发
频率控制别压垮目标服务器
登录数据不要采集非公开数据
# 尊重 robots.txt
from urllib.robotparser import RobotFileParser

rp = RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
can = rp.can_fetch("my-bot/1.0", "https://example.com/private/")
print("允许抓取?", can)

提醒:技术可行 ≠ 合法合规。批量抓取应优先考虑网站是否提供官方 API。


2. HTTP 请求:requests 与 httpx

2.1 requests 基础

import requests

session = requests.Session()          # 复用连接 + 保持 cookie
session.headers.update({"User-Agent": "my-scraper/1.0"})

resp = session.get("https://example.com/list?page=1", timeout=10)
print(resp.status_code, resp.url)
html = resp.text                      # 默认按 header 解码
print(resp.encoding)                  # 可能是 ISO-8859-1,需要处理中文
resp.encoding = "utf-8"               # 修正编码

2.2 httpx(同步/异步双模式)

import httpx

# 同步
with httpx.Client(headers={"User-Agent": "x"}, timeout=10) as client:
    r = client.get("https://example.com")

# 异步(高并发)
import asyncio

async def main():
    async with httpx.AsyncClient(timeout=10) as client:
        tasks = [client.get(f"https://example.com/item/{i}") for i in range(50)]
        results = await asyncio.gather(*tasks)
        for r in results:
            if r.status_code == 200:
                process(r.text)

asyncio.run(main())

2.3 处理分页与循环

def crawl_pages(base_url, max_pages=20):
    session = requests.Session()
    for page in range(1, max_pages + 1):
        resp = session.get(f"{base_url}?page={page}", timeout=10)
        if resp.status_code != 200:
            break
        yield parse_page(resp.text)
        time.sleep(0.5)   # 限速,别打太狠

3. 解析 HTML:BeautifulSoup 与 lxml

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "lxml")   # lxml 快;html.parser 更稳

# 提取标题
print(soup.title.text)

# 找单个元素
title = soup.find("h1")
# 按 class
items = soup.find_all("div", class_="product-item")
# 按属性
links = soup.find_all("a", href=True)
# CSS 选择器(更强大)
cards = soup.select("div.card > a.title")

for item in items:
    name = item.select_one("h2").text.strip()
    price = item.select_one(".price").text.strip()
    link = item.select_one("a")["href"]
    print(name, price, link)

lxml 直接解析(更快):

from lxml import html as lxml_html

tree = lxml_html.fromstring(html)
# XPath
prices = tree.xpath("//div[@class='price']/text()")

4. XPath 与 CSS 选择器

目标CSSXPath
id 选择#price//*[@id='price']
class.card//*[contains(@class,'card')]
子元素div > a//div/a
属性a[href^='http']//a[starts-with(@href,'http')]
文本—//h1/text()
第 n 个li:nth-child(2)//li[2]
# XPath 实战:提取所有商品链接与标题
links = tree.xpath("//div[@class='product']//a[@class='title']")
for a in links:
    text = a.text_content().strip()
    href = a.get("href")
    print(text, href)

调试选择器:浏览器 F12 → 复制 selector / XPath,再在 Python 里验证。


5. JSON API 与异步抓取

很多「网页」背后是 JSON API,直接调 API 更高效:

import httpx, asyncio, json

async def crawl_api():
    async with httpx.AsyncClient() as client:
        # 分页 JSON API
        all_items = []
        for page in range(1, 5):
            r = await client.get(
                "https://example.com/api/items",
                params={"page": page, "limit": 100},
            )
            data = r.json()
            all_items.extend(data["items"])
            if page >= data["total_pages"]:
                break

        # 并发抓详情
        async def fetch_detail(item_id):
            r = await client.get(f"https://example.com/api/items/{item_id}")
            return r.json()

        details = await asyncio.gather(
            *(fetch_detail(i["id"]) for i in all_items[:20])
        )
        return details

data = asyncio.run(crawl_api())
print(json.dumps(data, ensure_ascii=False, indent=2)[:500])

优点:请求少、数据干净、不受 HTML 改版影响。优先找 API。


6. Scrapy 框架

大规模爬虫用 Scrapy:自带调度、去重、管道、并发、中间件。

# items.py
import scrapy

class ProductItem(scrapy.Item):
    name = scrapy.Field()
    price = scrapy.Field()
    url = scrapy.Field()
# spiders/products.py
import scrapy

class ProductSpider(scrapy.Spider):
    name = "products"
    start_urls = ["https://example.com/products"]

    def parse(self, response):
        for card in response.css("div.product"):
            item = ProductItem()
            item["name"] = card.css("h2::text").get()
            item["price"] = card.css(".price::text").get()
            item["url"] = card.css("a::attr(href)").get()
            yield item

        # 自动翻页
        next_page = response.css("a.next::attr(href)").get()
        if next_page:
            yield response.follow(next_page, self.parse)
# 运行
scrapy crawl products -O products.json

Scrapy 核心优势:

特性说明
去重内置 URL 去重
并发自动调度异步请求
管道数据清洗/存储流水线
中间件代理、限速、UA 轮换
断点续爬JOBDIR

7. Playwright 浏览器自动化

需要执行 JavaScript 的页面(SPA、动态渲染)用 Playwright:

pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()

    page.goto("https://example.com/dashboard", wait_until="networkidle")
    page.wait_for_selector(".data-table")      # 等元素渲染

    # 交互:点击、输入、滚动
    page.click("#load-more")
    page.fill("#search", "python")
    page.wait_for_timeout(1000)

    # 提取
    rows = page.query_selector_all("tr")
    for row in rows[:5]:
        print(row.inner_text())

    # 截图 / PDF(便于调试)
    page.screenshot(path="debug.png")
    browser.close()

异步版 + 并发浏览器:

from playwright.async_api import async_playwright
import asyncio

async def snap(url):
    async with async_playwright() as p:
        b = await p.chromium.launch()
        pg = await b.new_page()
        await pg.goto(url)
        title = await pg.title()
        await b.close()
        return title

async def main():
    titles = await asyncio.gather(*[snap(u) for u in URLS])
    print(titles)

8. 反爬应对与限速

8.1 常见反爬手段与对策

反爬手段对策
UA 校验设置真实浏览器 UA
频率限制限速 + 指数退避重试
IP 封禁代理池轮换
JS 动态渲染Playwright
验证码人工兜底 / 服务降级
数据加密逆向或转 API
# 指数退避
def fetch_with_retry(url, max_retries=3):
    for i in range(max_retries):
        try:
            return requests.get(url, timeout=10)
        except requests.RequestException as e:
            wait = 2 ** i
            print(f"重试 {i+1},等待 {wait}s: {e}")
            time.sleep(wait)
    raise RuntimeError(f"最终失败: {url}")

8.2 限速设计(礼貌爬虫)

import time, random

def polite_fetch(url, min_delay=1.0, max_delay=3.0):
    time.sleep(random.uniform(min_delay, max_delay))  # 随机延时更像人
    return requests.get(url, timeout=10)

8.3 代理池

proxies = [
    "http://proxy1:8080",
    "http://proxy2:8080",
]
def get_proxy():
    return {"http": random.choice(proxies), "https": random.choice(proxies)}

态度:爬虫要保持「礼貌」——限速、带 UA、尊重 robots、失败退避。封禁往往源于抓太狠而非抓本身。


9. 数据清洗与存储

9.1 清洗管道

def clean_price(raw: str) -> float:
    """'¥1,299.00' → 1299.00"""
    import re
    s = re.sub(r"[^\d.]", "", raw)
    return float(s) if s else 0.0

def clean_text(raw: str) -> str:
    return " ".join(raw.split())   # 合并多余空白

# 结构化
records = [
    {"name": clean_text(name), "price": clean_price(price)}
    for name, price in raw_items
]

9.2 存储选项

数据量方案
小JSONL / CSV
中SQLite
大PostgreSQL / Parquet
需要变更追踪追加 JSONL + 定期去重
import json, sqlite3

# JSONL 追加
with open("items.jsonl", "a", encoding="utf-8") as f:
    f.write(json.dumps(record, ensure_ascii=False) + "\n")

# SQLite
conn = sqlite3.connect("items.db")
conn.execute("CREATE TABLE IF NOT EXISTS items (name TEXT, price REAL)")
conn.execute("INSERT INTO items VALUES (?, ?)", (record["name"], record["price"]))
conn.commit()

10. 速查表与最佳实践

场景工具
简单静态页requests + BeautifulSoup
大规模并发httpx 异步 / Scrapy
JSON APIrequests / httpx 直接调
JS 动态页Playwright
快速实验requests + lxml
数据量大Scrapy + PostgreSQL

最佳实践清单:

  1. 先看 robots.txt + 找官方 API。
  2. 用 Session/Client 复用连接。
  3. 设置超时、限速、退避重试。
  4. 统一数据 schema,清洗函数复用。
  5. 抓取结果落盘(JSONL),断点续抓。
  6. 出错要可恢复,别一崩全丢。
  7. 遵守目标网站 ToS,个人使用为主。

一句话记忆:爬虫四件套——requests 发请求、BeautifulSoup/lxml 解析、Scrapy 管规模、Playwright 过动态;永远先问「有没有 API、让不让抓」。

延伸阅读

爬虫是数据工程的起点,也是与 Web 生态互动的窗口。会抓只是第一步,会「干净地抓、合规地抓、可持续地抓」才是工程能力。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「python」更多文章

  1. Python 微服务架构:从单体拆分到服务治理
  2. Python 库与 API 设计:从包结构到向后兼容
  3. Python C 扩展与 FFI:ctypes、cffi、Cython 与 PyO3