实时通信最棘手的地方在于:用户说「卡」,但你无法复现。网络环境千差万别,问题往往只在特定运营商、特定机型、特定时段出现。没有一套可量化、可回溯的质量监控体系,优化就只能靠猜。
本文从指标体系讲起,覆盖数据采集、MOS 估算、埋点上报与告警,给出一套可落地的 QoE 监控方案。读完你应该能回答「我们的通话质量到底怎么样」这个问题,而不是依赖个别用户的反馈。
一、QoE 指标体系
QoE(Quality of Experience)是用户主观感受的量化。它由若干客观指标组合而成,每个指标反映不同维度:
| 指标 | 含义 | 目标值 | 采集方式 |
|---|---|---|---|
| 首帧时间 | 从发起到看到第一帧 | < 1s | 时间戳差值 |
| 卡顿率 | 卡顿时长占通话时长比 | < 1% | 播放事件 |
| 端到端延迟 | 语音往返延迟 | < 200ms | RTP 时间戳 + RTCP |
| 丢包率 | 丢包数占发送数比 | < 2% | getStats |
| MOS | 语音质量主观评分 | > 4.0 | E-model 估算 |
| 连接成功率 | 成功建连占比 | > 99% | 信令埋点 |
关键原则是:指标必须可分解到「设备、网络、地域、版本」等维度,否则发现了问题也无法定位原因。
二、用 getStats 采集
RTCPeerConnection.getStats() 是数据的源头,它返回一组 RTCStatsReport,涵盖 RTP、ICE、DTLS 各层统计。
async function collectStats(pc, label) {
const stats = await pc.getStats();
const result = { label, timestamp: Date.now() };
stats.forEach((report) => {
if (report.type === "inbound-rtp" && report.kind === "video") {
result.video = {
packetsReceived: report.packetsReceived,
packetsLost: report.packetsLost,
jitter: report.jitter,
framesDecoded: report.framesDecoded,
framesDropped: report.framesDropped,
freezeCount: report.freezeCount,
totalFreezesDuration: report.totalFreezesDuration,
frameWidth: report.frameWidth,
frameHeight: report.frameHeight,
};
}
if (report.type === "outbound-rtp" && report.kind === "video") {
result.outVideo = {
bytesSent: report.bytesSent,
framesEncoded: report.framesEncoded,
qualityLimitationReason: report.qualityLimitationReason,
};
}
if (report.type === "candidate-pair" && report.state === "succeeded") {
result.network = {
rtt: report.currentRoundTripTime,
availableOutgoingBitrate: report.availableOutgoingBitrate,
};
}
});
return result;
}
// 每秒采集一次
const samples = [];
const timer = setInterval(async () => {
samples.push(await collectStats(pc, "call"));
}, 1000);
qualityLimitationReason 是最有价值的字段之一,它直接告诉你视频质量受限的原因:bandwidth(带宽不足)、cpu(编码器过载)、other。看到它就能立刻判断瓶颈在编码侧还是网络侧。
三、MOS 估算
3.1 E-model 原理
MOS(Mean Opinion Score)原本是人工主观打分,范围 1 到 5。工程上用 E-model(ITU-T G.107)从客观指标估算:
function estimateMos({ rttMs, jitterMs, lossRate }) {
// 有效设备损伤因子 Ie
const ie = 10 + 25 * Math.log10(1 + lossRate * 100);
// 延迟损伤因子 Id
const id =
0.024 * rttMs +
0.11 * (rttMs - 177.3) * (rttMs > 177.3 ? 1 : 0) +
20 * (jitterMs > 0 ? Math.log10(1 + jitterMs) : 0);
// 基础质量 R 值
const r = 93.2 - ie - id;
// R 值映射到 MOS
if (r < 0) return 1;
if (r > 100) return 4.5;
return 1 + 0.035 * r + r * (r - 60) * (100 - r) * 7e-6;
}
console.log(estimateMos({ rttMs: 80, jitterMs: 10, lossRate: 0.01 })); // 约 4.2
console.log(estimateMos({ rttMs: 300, jitterMs: 50, lossRate: 0.08 })); // 约 2.6
MOS 的价值在于把多维指标压缩成一个可对比的分数,便于跨场景比较。但要注意它只是估算,不能替代真实的主观测试。
3.2 指标的分位统计
平均值会掩盖长尾问题。一个通话质量「平均 MOS 4.2」的系统,可能仍有 10% 的用户 MOS 低于 3。因此必须看分位数:
| 统计量 | 用途 |
|---|---|
| P50 | 反映典型体验 |
| P90 | 反映较差体验 |
| P99 | 反映极端劣化 |
| 低于阈值占比 | 反映问题用户比例 |
工程上通常把「MOS 低于 3.0 的通话占比」作为核心健康指标,目标控制在 5% 以内。
四、卡顿与首帧时间
4.1 卡顿率
卡顿率是最贴近用户感知的指标。它通过监听播放器的 waiting 与 playing 事件,或读取 totalFreezesDuration 计算:
class FreezeTracker {
constructor() {
this.freezeCount = 0;
this.freezeDuration = 0;
this.callStart = Date.now();
this.freezeStart = null;
}
onWaiting() {
if (this.freezeStart === null) this.freezeStart = Date.now();
}
onPlaying() {
if (this.freezeStart !== null) {
this.freezeCount++;
this.freezeDuration += Date.now() - this.freezeStart;
this.freezeStart = null;
}
}
rate() {
const total = Date.now() - this.callStart;
return total > 0 ? this.freezeDuration / total : 0;
}
}
注意卡顿率要区分「音视频整体卡顿」与「仅视频卡顿」。音频卡顿的感知严重得多,应单独统计。
4.2 首帧时间
首帧时间(Time to First Frame)从用户点击发起,到远端画面首次渲染:
const t0 = performance.now();
pc.ontrack = (event) => {
const video = document.querySelector("#remote");
video.srcObject = event.streams[0];
video.addEventListener("loadeddata", () => {
const ttff = performance.now() - t0;
report("ttff", ttff);
}, { once: true });
};
首帧时间受信令、ICE 收集、DTLS 握手、解码器初始化等多个环节影响,是端到端体验的综合体现。Trickle ICE 与预置 TURN 候选能显著降低这一指标。
五、埋点与上报
5.1 采样与聚合
逐秒上报原始数据会造成巨大流量,正确做法是客户端聚合后再上报:
class MetricsAggregator {
constructor() {
this.samples = [];
}
push(sample) {
this.samples.push(sample);
}
// 通话结束时聚合上报
flush(callId) {
if (this.samples.length === 0) return;
const loss = this.samples.map((s) => s.video?.packetsLost || 0);
const payload = {
callId,
duration: this.samples.length,
avgLoss: average(loss),
maxLoss: Math.max(...loss),
avgRtt: average(this.samples.map((s) => s.network?.rtt || 0)),
qualityLimitation: this.samples.map((s) => s.outVideo?.qualityLimitationReason),
};
navigator.sendBeacon("/metrics", JSON.stringify(payload));
}
}
function average(arr) {
return arr.length ? arr.reduce((a, b) => a + b, 0) / arr.length : 0;
}
navigator.sendBeacon 在页面卸载时也能可靠发送,适合通话结束时的最终上报。
5.2 上报的维度
每条记录必须携带足够的维度标签,否则数据无法用于定位:
| 维度 | 示例 | 用途 |
|---|---|---|
| 设备 | iPhone 15 / Chrome 120 | 机型兼容问题 |
| 网络类型 | Wi-Fi / 4G / 5G | 网络差异 |
| 地域 | 运营商 + 省份 | 区域性问题 |
| 应用版本 | v1.2.3 | 版本回归 |
| 角色 | 发布者 / 订阅者 | 定位方向 |
六、可视化与告警
采集到的数据要能被人看懂才有价值。推荐搭建的看板包括:
- 实时大盘:当前通话数、连接成功率、平均 MOS、卡顿率。
- 趋势图:各指标随时间变化,便于发现劣化拐点。
- 维度下钻:按设备、地域、版本分解,定位问题来源。
- 单通话回放:针对投诉的通话,还原其指标曲线。
告警规则应基于分位数而非平均值,例如「P90 卡顿率连续 5 分钟超过 3%」触发告警。
七、连接成功率与信令埋点
7.1 为什么单独统计
音视频指标只覆盖「已经连上」的通话。连接失败的用户根本产生不了媒体指标,却往往是体验最差的一群人。因此必须单独统计连接成功率,并把它拆解到 ICE 的各个阶段。
const connectMetrics = {
offerCreated: false,
answerReceived: false,
iceConnected: false,
dtlsConnected: false,
mediaFlowing: false,
failedAt: null,
};
pc.oniceconnectionstatechange = () => {
if (pc.iceConnectionState === "connected") {
connectMetrics.iceConnected = true;
}
if (pc.iceConnectionState === "failed") {
connectMetrics.failedAt = "ice";
reportConnectFailure(connectMetrics);
}
};
pc.onconnectionstatechange = () => {
if (pc.connectionState === "connected") {
connectMetrics.dtlsConnected = true;
}
if (pc.connectionState === "failed") {
connectMetrics.failedAt = connectMetrics.iceConnected ? "dtls" : "ice";
reportConnectFailure(connectMetrics);
}
};
7.2 失败原因的分类
把失败原因归类,才能针对性优化:
| 失败阶段 | 典型原因 | 优化方向 |
|---|---|---|
| 信令未建立 | WebSocket 连不上 | 信令就近部署、降级到轮询 |
| 无候选 | STUN 不可达 | 部署多地域 STUN |
| ICE 失败 | 无可用候选对 | 部署 TURN |
| DTLS 失败 | 证书或指纹问题 | 检查 SDP 改写 |
| 无媒体 | 编解码不兼容 | 放宽协商范围 |
连接成功率的目标通常在 99% 以上,且必须按网络类型与地域分解,因为某些运营商或区域的失败率可能显著偏高。
八、A/B 与回归验证
任何优化上线前都应有数据支撑。做法是灰度一部分用户到新版本,对比两组的 QoE 指标:
// 上报时附带实验分组
const group = hashUserId(userId) % 100 < 10 ? "treatment" : "control";
report("qoe", { group, mos, freezeRate, ttff });
对比时要关注统计显著性,避免小样本波动误导。核心指标改善而次要指标不劣化,才算一次成功的优化。
八、常见坑清单
- 只看平均值不看分位数,掩盖了长尾用户的问题。
- 逐秒上报原始数据,流量与存储成本失控。
- 上报缺少维度标签,发现问题却无法定位。
- MOS 公式参数未按场景校准,估算值与真实感受偏差大。
- 卡顿率不区分音视频,把视频轻微卡顿与音频断续混为一谈。
- 忽略
qualityLimitationReason,无法区分是带宽问题还是 CPU 问题。 - 页面卸载时用普通请求上报,数据大量丢失。
- 告警基于瞬时值,网络抖动导致误报频繁。
小结
QoE 监控的价值在于把「用户觉得卡」转化为可量化、可定位、可回归的指标。核心是把首帧时间、卡顿率、MOS、丢包率这些指标采集全,并携带足够的维度标签。getStats 是数据源头,qualityLimitationReason 是判断瓶颈方向的关键字段,分位数统计是发现长尾问题的前提。指标体系的建设应与 弱网对抗与自适应码率
的优化形成闭环:先量化,再优化,再验证。没有监控的优化是盲目的,而监控体系本身也需要随业务演进持续打磨。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。