《Python编程实战》3.3 指标、健康检查与告警接入

日志记录离散事件,指标记录系统整体健康度。本节用 prometheus-client 0.26.0 暴露 Counter/Gauge/Histogram,用中间件埋点并规避标签基数陷阱,设计 liveness/readiness 健康检查,再接入 Prometheus 告警规则。

本节目标:用 prometheus-client 0.26.0 暴露指标端点,用中间件埋点并规避标签基数陷阱,设计区分 liveness/readiness 的健康检查,最后接入 Prometheus 告警规则。
适用版本:Python 3.12+(实测 3.14.6);prometheus-client 0.26.0;fastapi 0.143.0

3.3 指标、健康检查与告警接入

前两节我们解决了「配置从哪来」和「运行中发生了什么」。但日志是一串离散事件,想知道「此刻服务到底健不健康」,靠翻日志太慢。本节补上可观测性的第三块拼图——指标,并把它接到健康检查与告警上。

3.3.1 可观测性三支柱与指标的定位

支柱回答的问题数据形态成本
日志单个事件为什么发生离散、带上下文高(量大)
指标系统整体趋势如何聚合、时间序列低(预聚合)
追踪一次请求慢在哪一段有向调用链中

三者的分工是:指标负责「发现异常」,日志和追踪负责「定位原因」。指标的价值在于它的低成本和可聚合——每 15 秒采集一次,几百个时间序列就能画出整个服务的请求量、错误率、延迟分布。Prometheus 是事实标准,Python 侧的官方客户端就是 prometheus-client。

3.3.2 四种指标类型

prometheus-client 提供四种类型,选错类型会让后续聚合变得别扭:

类型语义典型用途
Counter只增不减的累计值请求总数、错误总数
Gauge可增可减的瞬时值队列深度、连接数、内存
Histogram分桶统计分布请求延迟、响应体大小
Summary客户端分位数精确分位数(少用,难聚合)

Counter 名字必须以 _total 结尾是社区约定;Histogram 用一组 le(小于等于)桶来近似分布,能算出 P99 这类分位数。下面把四种都跑一遍:

from prometheus_client import Counter, Gauge, Histogram, Summary, generate_latest, CollectorRegistry

reg = CollectorRegistry()
requests = Counter("http_requests_total", "HTTP 请求总数", ["method", "path", "status"], registry=reg)
inflight = Gauge("http_inflight_requests", "当前处理中的请求数", registry=reg)
latency = Histogram("http_request_duration_seconds", "请求耗时",
                    buckets=[0.05, 0.1, 0.25, 0.5, 1.0, 2.5], registry=reg)

inflight.inc()
requests.labels(method="GET", path="/api/orders", status="200").inc()
requests.labels(method="GET", path="/api/orders", status="200").inc()
requests.labels(method="POST", path="/api/orders", status="500").inc()
latency.observe(0.083)
latency.observe(0.42)
inflight.dec()

print(generate_latest(reg).decode())
# HELP http_requests_total HTTP 请求总数
# TYPE http_requests_total counter
http_requests_total{method="GET",path="/api/orders",status="200"} 2.0
http_requests_total{method="POST",path="/api/orders",status="500"} 1.0
# HELP http_inflight_requests 当前处理中的请求数
# TYPE http_inflight_requests gauge
http_inflight_requests 0.0
# HELP http_request_duration_seconds 请求耗时
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{le="0.05"} 0.0
http_request_duration_seconds_bucket{le="0.1"} 1.0
http_request_duration_seconds_bucket{le="0.5"} 2.0
http_request_duration_seconds_bucket{le="+Inf"} 2.0
http_request_duration_seconds_count 2.0
http_request_duration_seconds_sum 0.503

(为简洁省略了自动生成的 _created 系列与部分中间桶。)

这就是 Prometheus 的文本 exposition 格式。注意 Histogram 自动生成了三组序列:_bucket(各桶计数)、_count(总次数)、_sum(总和)。le="0.1" 桶里是 1.0,说明两次请求里只有一次耗时 ≤ 0.1 秒;_count 为 2、_sum 为 0.503,两者相除即平均耗时。分位数不直接给出,而是由 Prometheus 用 histogram_quantile() 在查询时算出来——这正是 Histogram 能被跨实例聚合的原因。

3.3.3 暴露 /metrics 端点

指标要被 Prometheus 抓取,就得暴露一个 HTTP 端点。generate_latest() 返回文本格式的字节串,配上正确的 Content-Type 即可:

from fastapi import FastAPI, Response
from prometheus_client import Counter, generate_latest, CONTENT_TYPE_LATEST

app = FastAPI()
REQUEST_COUNT = Counter("app_requests_total", "请求总数", ["path"])

@app.get("/metrics")
def metrics():
    REQUEST_COUNT.labels(path="/metrics").inc()
    return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)

CONTENT_TYPE_LATEST 是 text/plain; version=1.0.0; charset=utf-8,Prometheus 靠它识别格式。生产上要注意两点:/metrics 端点别暴露到公网(会泄露内部结构与流量),以及多进程部署要用 prometheus_client.multiprocess 模式,否则每个 worker 各自计数、抓到的数据不完整。

3.3.4 中间件埋点与标签基数陷阱

指标要在中间件里统一埋点,而不是每个路由手写。但这里藏着一个高频事故——用真实路径当标签:

@app.middleware("http")
async def observe(request: Request, call_next):
    t0 = time.perf_counter()
    response = await call_next(request)
    # 直接用 request.url.path 当标签
    REQ.labels(request.method, request.url.path, str(response.status_code)).inc()
    return response
http_requests_total{method="GET",path="/api/items/1",status="200"} 1.0
http_requests_total{method="GET",path="/api/items/2",status="200"} 1.0

每次请求的 item_id 不同,就产生一条全新的时间序列。用户 ID、订单号这类高基数维度一旦进标签,时间序列会爆炸式增长,Prometheus 内存被拖垮。正确做法是用路由模板(/api/items/{item_id})而不是真实路径:

@app.middleware("http")
async def observe(request: Request, call_next):
    t0 = time.perf_counter()
    response = await call_next(request)
    # 用路由模板,避免标签基数爆炸
    route = request.scope.get("route")
    template = getattr(route, "path", request.url.path)
    LAT.labels(template).observe(time.perf_counter() - t0)
    REQ.labels(request.method, template, str(response.status_code)).inc()
    return response
http_requests_total{method="GET",route="/api/items/{item_id}",status="200"} 3.0
http_request_duration_seconds_count{route="/api/items/{item_id}"} 3.0

三次不同 item_id 的请求被聚合到同一条序列上。request.scope["route"] 是 Starlette 在路由匹配后写入的,拿到的就是模板字符串。标签只放基数可控的维度(方法、路由模板、状态码),高基数的 ID 交给日志和追踪去承载。

3.3.5 Gauge 的实时值与 set_function

有些指标不是「累加」而是「此刻是多少」,比如队列深度、活跃连接数。Gauge 可以直接 set,也可以用 set_function 在每次被采集时现算,省去维护内部状态:

import psutil
from prometheus_client import Gauge, generate_latest, CollectorRegistry

reg = CollectorRegistry()
rss = Gauge("process_resident_memory_bytes", "常驻内存", registry=reg)
rss.set(psutil.Process().memory_info().rss)

queue = Gauge("task_queue_depth", "队列深度", registry=reg)
pending = [1, 2, 3, 4]
queue.set_function(lambda: len(pending))   # 采集时现算

print(generate_latest(reg).decode())
process_resident_memory_bytes 3.0212096e+07
task_queue_depth 4.0

(省略了 # HELP / # TYPE 注释行。)

set_function 的好处是「永远读到最新值」,避免埋点代码忘了更新 Gauge 导致读数失真。但要注意回调要轻量——它在每次 /metrics 被抓取时同步执行,里若做重活会拖慢采集。

3.3.6 健康检查:liveness 与 readiness

健康检查是编排平台(K8s、负载均衡)判断实例状态的手段,但必须区分两种语义,混用会引发级联故障:

探针问题失败后果该检查什么
liveness(存活)进程还活着吗重启容器只检查进程自身,别查依赖
readiness(就绪)能接流量了吗摘除流量检查依赖(DB、缓存)是否就绪

核心原则:liveness 绝不能检查下游依赖。如果数据库抖动导致 liveness 失败,K8s 会把所有健康实例一起重启,反而放大故障。下面用 FastAPI 实现两个端点:

from fastapi import FastAPI, Response

app = FastAPI()
_ready = False

@app.get("/healthz")           # liveness:进程活着就 200
def healthz():
    return {"status": "ok"}

@app.get("/readyz")            # readiness:依赖就绪才 200
def readyz(response: Response):
    if not _ready:
        response.status_code = 503
        return {"status": "not_ready", "checks": {"db": "down"}}
    return {"status": "ready", "checks": {"db": "up"}}

用 TestClient 验证:

from fastapi.testclient import TestClient

client = TestClient(app)
print("GET /healthz ->", client.get("/healthz").status_code, client.get("/healthz").json())
r = client.get("/readyz")
print("GET /readyz  ->", r.status_code, r.json())
_ready = True                      # 模拟依赖就绪
print("GET /readyz  ->", client.get("/readyz").status_code, client.get("/readyz").json())
GET /healthz -> 200 {'status': 'ok'}
GET /readyz  -> 503 {'status': 'not_ready', 'checks': {'db': 'down'}}
GET /readyz  -> 200 {'status': 'ready', 'checks': {'db': 'up'}}

/healthz 永远返回 200(只要进程还能响应,就说明它活着);/readyz 在依赖未就绪时返回 503——503 让负载均衡把它摘掉,而不是重启它。(本机运行 TestClient 会打出一条 httpx 相关的 StarletteDeprecationWarning,属环境提示,不影响结果。)readiness 返回体里带上 checks 明细,排障时一眼能看出是哪个依赖没起来。

3.3.7 让 Prometheus 抓取

指标暴露出来后,需要在 Prometheus 侧配置抓取目标。这是运维配置,本机未实测,仅示意:

scrape_configs:
  - job_name: orders-api
    metrics_path: /metrics
    scrape_interval: 15s
    static_configs:
      - targets: ["orders-api:8000"]

K8s 环境通常用 ServiceMonitor(Prometheus Operator)或 Pod 注解自动发现,无需手写 static_configs。

3.3.8 告警接入:症状告警优先

指标最终要变成告警。核心原则是对症状(用户感知)告警,而不是对原因(内部指标)告警——「5xx 比例升高」是症状,「CPU 高」是原因,前者才值得半夜叫醒人。下面是基于上面埋点的告警规则(Prometheus 配置,本机未实测,仅示意):

groups:
  - name: orders-api-slo
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[5m]))
          / sum(rate(http_requests_total[5m])) > 0.05
        for: 10m
        labels:
          severity: critical
        annotations:
          summary: "5xx 错误率超过 5%"
          runbook: "https://runbook.example.com/high-error-rate"
      - alert: HighLatencyP99
        expr: |
          histogram_quantile(0.99,
            sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
          ) > 1.0
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "P99 延迟超过 1s"

几个要点:rate(...[5m]) 算的是每秒增长率,sum(rate(...)) 把各标签聚合;for: 10m 要求异常持续 10 分钟才触发,避免毛刺误报;histogram_quantile(0.99, ...) 从桶算出 P99——这正是 3.3.2 里 Histogram 埋点的回报。告警里附 runbook 链接,是让「被叫醒的人」知道第一步该做什么。

告警设计反例正例
告警对象CPU > 80%5xx 比例 > 5%
阈值拍脑袋定基于 SLO 错误预算
持续时长一有波动就告警for 持续 10 分钟
可操作性只报「异常」附 runbook 与影响面

3.3.9 三件套的协作

指标、日志、追踪不是三套独立系统,而是同一套可观测性能力的三个视角。一个典型的排障闭环是:

  1. 指标告警触发——HighErrorRate 说明 5xx 涨了。
  2. 看日志——按 request_id/trace_id 过滤,找到具体报错事件。
  3. 看追踪——点开 trace_id,看这次请求慢/错在哪个 span。

这也是为什么上一节要让 trace_id 进日志:只有 ID 一致,三个视角才能互相跳转。想更系统地了解可观测性工程,可延伸阅读 Python 调试与日志工程 。

小结

  • 可观测性三支柱分工明确:指标负责「发现异常」,日志与追踪负责「定位原因」。
  • prometheus-client 四类型:Counter(只增)、Gauge(可增可减)、Histogram(分桶分布)、Summary(客户端分位数,少用)。
  • generate_latest() 输出文本 exposition 格式,配 CONTENT_TYPE_LATEST 暴露 /metrics,该端点勿对公网开放。
  • 中间件埋点要用路由模板而非真实路径当标签,高基数 ID 进标签会导致时间序列爆炸。
  • 健康检查必须区分 liveness(只查进程,失败即重启)与 readiness(查依赖,失败即摘流量),liveness 绝不能检查下游。
  • 告警对症状而非原因,配合 rate()、histogram_quantile()、for 持续时长与 runbook 链接。

到这里,「工程基建」部分就完整了:配置、日志、指标三件套齐备,服务「跑起来」的底座搭好。下一章我们进入测试工程化,用 pytest 把质量门禁固化进流程。

阅读导航:上一节:结构化日志与链路追踪 · 下一节:pytest 工程化:fixture 分层与插件 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「python」更多文章

  1. 《Python高级编程》目录
  2. 《Python高级编程》11.3 PEP 流程与版本迁移策略
  3. 《Python高级编程》11.2 嵌入式与自由线程运行时