26. MongoDB 监控与可观测性实践

MongoDB 指标体系(currentOp/serverStatus/dbStats)、慢查询与 Profiling(system.profile)、Oplog 监控、Prometheus + mongodb-exporter + Grafana、告警规则设计与连接数分析

MongoDB 不是黑盒,但它的可观测性需要刻意建设。与关系型数据库成熟的监控体系相比,MongoDB 的指标散布在 serverStatus、currentOp、dbStats 与副本集状态之中,稍不留意就会漏掉关键的"先行指标"——例如 oplog 窗口缩小、应用线程被迫驱逐缓存、连接数逼近上限。本文将从 MongoDB 原生指标入手,再到 Prometheus + Grafana 的采集可视化,最后给出经过实战检验的告警规则设计,帮助团队建立从"出了事才看日志"到"指标先行"的监控体系。

1. MongoDB 指标体系总览

MongoDB 的运维指标可以划分为六类:操作吞吐、延迟、连接、内存与缓存、复制、持久性。绝大多数都能从三个入口获取:db.serverStatus()、db.currentOp()、db.stats()。

// 服务端全局状态(核心入口)
db.serverStatus()
// { "host": "...", "version": "7.0.x", "uptime": 86400,
//   "connections": {...}, "opcounters": {...}, "network": {...},
//   "mem": {...}, "wiredTiger": {...}, "repl": {...}, ... }

// 会话级:查看当前正在执行的操作
db.currentOp({ active: true, $or: [{ op: "query" }, { op: "command" }] })

db.serverStatus() 返回数百个字段,实践中应聚焦以下核心指标:

分类关键字段反映的问题
吞吐opcounters.insert/query/update/delete各操作 QPS 趋势
延迟opLatencies读写命令延迟 p95/p99
连接connections.current/available连接池健康度
内存mem.resident/virtual, mem.mapped常驻内存与映射
缓存wiredTiger.cache.*命中率与驱逐压力
复制repl.setName, repl.state副本集角色与健康
存储dbStats.size/storageSize数据量与碎片
// 简洁的"一分钟体检"
db.serverStatus().opcounters
db.serverStatus().connections
db.serverStatus().mem
db.serverStatus().wiredTiger.cache["dirty percentage in the cache"]
db.getReplicationInfo()   // oplog 窗口

2. 慢查询与 Profiling

慢查询是可观测性的第一优先。MongoDB 的 Profiler 会把超过阈值(slowms)的操作写入 system.profile 固定大小集合。

// 开启慢查询收集(只记录 > 100ms 的操作)
db.setProfilingLevel(1, { slowms: 100 })

// 查看当前 profiling 状态
db.getProfilingStatus()
// { was: 1, slowms: 100, sampleRate: 1 }

// 查询最近的慢查询
db.system.profile.find().sort({ ts: -1 }).limit(10).toArray()

system.profile 记录了操作的执行计划、锁等待与读写字节数,是定位慢查询根因的第一手数据:

// 按集合聚合慢查询:平均耗时、次数、最大耗时
db.system.profile.aggregate([
  { $match: { op: "query", millis: { $gt: 100 } } },
  { $group: {
      _id: "$ns",
      avgMillis: { $avg: "$millis" },
      maxMillis: { $max: "$millis" },
      count: { $sum: 1 }
  }},
  { $sort: { avgMillis: -1 } }
])

// 查看某条慢查询是否走了 COLLSCAN
db.system.profile.find(
  { "ns": "shop.orders", "planSummary": /COLLSCAN/ },
  { "millis": 1, "query": 1, "planSummary": 1 }
).sort({ millis: -1 }).limit(5)
Profiling 级别含义生产建议
0关闭默认
1仅记录慢查询(slowms)推荐长期开启,slowms 调至 100~300
2记录所有操作仅短时排查使用

注意:system.profile 是 capped 集合(默认 1MB),高负载下旧记录会被覆盖。生产环境建议将 slowms 提高到 100ms+ 并配合 sampleRate(如 0.5)抽样,降低 profiler 自身的写入开销。长期采集应通过 Prometheus 落库或定期导出归档。

3. Oplog 监控:滞后又可写

oplog 是复制集的心跳。两个最关键的指标是复制滞后(replication lag)与oplog 窗口(oplog window)。

  • 复制滞后:Secondary 落后 Primary 的时间,反映同步压力
  • oplog 窗口:oplog 中现存数据可覆盖的时间范围,反映"还能回放到多远"
// 副本集复制概况
rs.printReplicationInfo()
// configured oplog size:   1879MB
// log length start to end: 86400 (24 hours)

// 各 Secondary 的同步位置与滞后
rs.printSecondaryReplicationInfo()
// source: 172.17.0.2:27017
//   syncedTo: Thu Sep 27 2026 12:00:00
//   0 secs (0 hrs) behind the primary

// 直接读取 oplog 首尾时间戳
db.getSiblingDB("local").oplog.rs.find().sort({ $natural: 1 }).limit(1).next().ts
db.getSiblingDB("local").oplog.rs.find().sort({ $natural: -1 }).limit(1).next().ts

oplog 窗口缩小的典型原因:写入速率高、oplogSizeMB 配置过小、某个 Secondary 长期离线导致它需要的窗口不断增长。窗口小于一定阈值(如 24 小时)会威胁 Point-in-Time 恢复与在线迁移能力。

// 增大 oplog 窗口(需要重启生效的经典做法)
// 1) 对副本集某节点停止服务
// 2) 以 --oplogSize MB 重启
// 或使用 replSetResizeOplog(4.4+ 可在线调整)
db.adminCommand({ replSetResizeOplog: 1, size: 4096 })  // 4096 MB
Oplog 指标命令健康阈值
复制滞后rs.printSecondaryReplicationInfo()< 10 秒
窗口时长rs.printReplicationInfo()> 24 小时
oplog 使用率db.getReplicationInfo()< 80%

4. Prometheus 采集:mongodb-exporter

生产环境推荐用 mongodb-exporter(Percona 维护)把 MongoDB 指标转换为 Prometheus 格式,再由 Grafana 可视化。exporter 通过 serverStatus、replSetGetStatus 等命令采集指标,也支持 profiling 采集。

# docker-compose.yml:mongodb-exporter 部署
services:
  mongodb-exporter:
    image: percona/mongodb_exporter:0.40
    command:
      - "--mongodb.uri=mongodb://monitor:pass@mongodb-primary:27017/admin?replicaSet=rs0"
      - "--discovering-mode"
      - "--compatible-mode"
      - "--collector.collection"
      - "--collector.topmetrics"
    ports:
      - "9216:9216"
    restart: unless-stopped
# 验证指标输出
curl -s http://localhost:9216/metrics | head -n 20
# mongodb_mongod_connections_current 120
# mongodb_mongod_connections_available 49999
# mongodb_mongod_global_lock_currentqueue_total{type="read"} 0
# mongodb_mongod_metrics_cursor_open_total{state="total"} 12

Prometheus 采集配置(prometheus.yml):

scrape_configs:
  - job_name: mongodb
    static_configs:
      - targets: ["mongodb-exporter:9216"]
exporter 指标对应 MongoDB 指标监控价值
mongodb_mongod_connections_currentconnections.current连接瓶颈预警
mongodb_mongod_opcounters_query_totalopcounters.query读吞吐
mongodb_mongod_wiredtiger_cache_dirty_byteswiredTiger.cache dirty写缓冲压力
mongodb_mongod_replset_member_replication_lagrepl lag同步健康
mongodb_mongod_op_latencies_commands_secondsopLatencies命令延迟
mongodb_mongod_db_collection_total_documentscollStats数据量增长

5. Grafana 仪表盘设计

Grafana 侧重点是"一眼看出异常"而不是"信息越多越好"。建议按时间线组织面板,从概览到钻取共三层。

第一层:集群总览

  • 连接数(当前/可用)
  • 各操作 QPS(opcounters)
  • 读写延迟 p95/p99
  • 复制滞后
  • oplog 窗口时长

第二层:引擎与资源

  • WiredTiger cache:命中率、dirty 百分比、驱逐页数
  • 内存:resident / mapped / virtual
  • 磁盘:storageSize、journal 目录占用
  • 慢查询计数(rate 形式)

第三层:业务钻取

  • 按集合的文档数与 storageSize
  • 按命名空间的 opcounters
  • 锁等待与 currentOp 活跃数
# 一个典型的 PromQL 示例:查询 5 分钟内慢查询速率
# rate(mongodb_mongod_slow_queries_total[5m])

建议:不要把数百个指标堆在一个面板。先建设"第 1 层 + 第 2 层"共 8~12 个核心面板,稳定运行两周后再根据事故复盘补充钻取面板。

6. 告警规则设计

告警规则设计的原则是"告警要能落地行动"。每个告警都应回答三个问题:谁负责处理、是否真的异常、如何定位。下面是一组经过生产验证的 Prometheus 告警规则。

# prometheus-rules.yml
groups:
  - name: mongodb_critical
    rules:
      - alert: MongoDBPrimaryDown
        expr: mongodb_mongod_replset_member_state{state="PRIMARY"} == 0
        for: 1m
        labels: { severity: critical }
        annotations:
          summary: "MongoDB 主节点不可用"

      - alert: MongoDBReplicationLagHigh
        expr: mongodb_mongod_replset_member_replication_lag > 10
        for: 2m
        labels: { severity: warning }
        annotations:
          summary: "复制滞后超过 10 秒"

      - alert: MongoDBOplogWindowTooSmall
        expr: time() - mongodb_mongod_replset_oplog_tail_timestamp > 86400
        for: 10m
        labels: { severity: critical }
        annotations:
          summary: "oplog 窗口小于 24 小时,恢复能力受限"

      - alert: MongoDBConnectionsExhausted
        expr: mongodb_mongod_connections_current / mongodb_mongod_connections_available > 0.8
        for: 5m
        labels: { severity: warning }
        annotations:
          summary: "连接数超过可用连接的 80%"

      - alert: MongoDBCacheEvictionPressure
        expr: rate(mongodb_mongod_wiredtiger_cache_pages_evicted_total[5m]) > 100
        for: 10m
        labels: { severity: warning }
        annotations:
          summary: "WiredTiger 驱逐压力增大,检查 cacheSizeGB"
告警级别响应时限示例
critical立即主节点不可用、oplog 窗口过小
warning15 分钟内复制滞后、连接数告急、驱逐压力
info当日处理慢查询比例升高、磁盘增长超预期

告警阈值校准:固定阈值容易误报。连接数告警应参考基线(例如当前/可用 > 0.8 且持续 5 分钟);慢查询告警用 rate 而不是绝对值,避免业务波动触发噪音。

7. 连接数与工作负载分析

连接数是最容易被忽视、又最常在高峰暴露的指标。MongoDB 每个连接消耗约 1MB 内存(连接上下文 + 线程栈),maxIncomingConnections 默认 65536,但应用连接池可能把请求堆积转化为连接风暴。

// 连接数实况
db.serverStatus().connections
// { current: 320, available: 65216, totalCreated: 12800, ... }

// 找出占用连接的来源(应用/监控/备份工具)
db.currentOp({ "connectionId": { $exists: true } }).inprog.forEach(op => {
  print(op.client, op.connectionId)
})

工作负载分析的关键问题:连接数暴涨是因为 QPS 上升,还是因为连接池配置错误(如 maxPoolSize 过大、空闲连接不回收)?把连接数、opcounters、延迟画在同一时间轴对比即可判断。

// 查看驱动上报的等待队列(若有 waitQueue 指标)
// 4.4+ 通过 mongo shell 观察 db.currentOp() 中等待锁的操作占比
db.currentOp({ waitForLock: true }).inprog.length

经验:连接数告急时优先检查是否有人在用"一次请求建一个连接"的反模式,或 mongod 的 maxIncomingConnections 小于应用连接池总和。扩容前先调配置,通常比加机器便宜得多。

8. 与 Atlas 监控对比

Atlas 托管集群内置了完整的监控体系,很多团队迁移到 Atlas 后反而失去了"自己搭监控"的能力,需要理解两者的定位差异。

能力自建 + Prometheus/GrafanaAtlas 内置监控
指标采集exporter 自建平台自动
慢查询分析Profiler + system.profileReal-Time Performance Panel
索引建议自行分析 $indexStatsPerformance Advisor 自动推荐
告警通道Prometheus + AlertmanagerEmail/Slack/PagerDuty
数据保留自定平台策略
定制自由度高中

Atlas 的 Real-Time Performance Panel 可以实时查看每个操作的计划摘要与延迟分布;Performance Advisor 会基于历史工作负载给出索引建议与分片建议。自建环境可用 mtools(mprof、mplotqueries、mlogfilter)模拟部分体验:

# mtools:解析慢日志并可视化
mplotqueries --type histogram mongod.log
mprof mongod.log | mplotqueries --type histogram --format PNG > profile.png
mlogfilter mongod.log --slow --threshold 100

对于缺乏专职 DBA 的团队,Atlas 的自动化能力(自动索引推荐、内置告警)能显著降低排障成本;而自建集群的监控自由度更高,可完全贴合业务定制。两者的共同底线是:指标必须先行、告警必须可行动、排障必须可回放。

延伸阅读

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「mongodb」更多文章

  1. 27. MongoDB 多租户与隔离架构设计
  2. 25. MongoDB WiredTiger 存储引擎深入
  3. 24. MongoDB 数据迁移与同步实战