《Go 语言运行时原理》6.2 GC 阶段、辅助标记与 pacing

把 gctrace 一行拆成四段时间与三段 CPU,再用 runtime/metrics 读出 STW 暂停分布(p50 18µs / p99 61µs)和 GC 的 CPU 去向;并钉到 go1.27.0 源码——setGCPhase 的阶段开关、gcAssistAlloc 的债务换算、gcControllerState 的 consMark 与 runway,最后给出读 GC 指标的顺序。

6.2 GC 阶段、辅助标记与 pacing

上一节讲了单次指针写入的屏障成本。这一节把尺度放大到一整轮 GC:一轮 GC 由哪些阶段组成、哪些阶段会 STW、标记工作由谁完成、以及运行时怎么决定「下一轮什么时候开始」。三色标记、写屏障与并发回收的原理在 /posts/golang/ 下已有专题文章讲清,本节不重复,只写 gctrace 与 runtime/metrics 的实测增量。

三个问题在工程上最常被问到:

  • gctrace 那一长串数字分别是什么?
  • 为什么有时 GC 会「卡住应用」——是 STW 变长了,还是辅助标记在偷 CPU?
  • 我改了 GOGC/GOMEMLIMIT,为什么 GC 次数变了但暂停没变?

本节用 gctrace 与 runtime/metrics 两组数据回答它们。

本节要回答:GC 的阶段、辅助标记与 pacing 各自在指标上长什么样?结论是:一轮 GC 只有两次 STW(开始时的 mark setup 与结束时的 mark termination),本机实测暂停 p50 约 18 µs、p99 约 61 µs;标记工作由 assist(分配方按债务补偿)、dedicated(专职 worker)、idle(空闲 P)三类共同完成;pacing 由 consMark(分配速率与扫描速率之比)与 runway 决定下一轮的触发点,所以调 GOGC 改的是「次数」,GOMEMLIMIT 改的是「上限」,两者都不直接缩短单次暂停。

6.2.1 实验一:把 gctrace 一行拆开

复现基线:

  • Go 工具链 go version go1.27.0 darwin/arm64(GOTOOLCHAIN=go1.27.0)
  • 机器:Apple M1 Pro,10 核,32 GiB;GOMAXPROCS 默认(10)
  • 负载:持续分配含指针对象,跑约 2 秒;用 GODEBUG=gctrace=1 输出

原始输出:

$ GOTOOLCHAIN=go1.27.0 GODEBUG=gctrace=1 ./prog
gc 1 @0.003s 3%: 0.041+1.8+0.013 ms clock, 0.41+0/1.8/0.036+0.13 ms cpu, 4->5->3 MB, 4 MB goal, 0 MB stacks, 0 MB globals, 10 P
gc 2 @0.006s 5%: 0.004+1.7+0.009 ms clock, 0.046+0/2.4/0.92+0.090 ms cpu, 6->6->4 MB, 7 MB goal, 0 MB stacks, 0 MB globals, 10 P
gc 3 @0.008s 7%: 0.008+2.4+0.018 ms clock, 0.086+0/3.6/2.0+0.18 ms cpu, 8->8->6 MB, 9 MB goal, 0 MB stacks, 0 MB globals, 10 P
gc 4 @0.012s 8%: 0.011+4.0+0.004 ms clock, 0.11+0.19/6.1/3.0+0.048 ms cpu, 13->13->10 MB, 13 MB goal, 0 MB stacks, 0 MB globals, 10 P
gc 5 @0.019s 9%: 0.029+6.2+0.012 ms clock, 0.29+0.67/9.3/4.6+0.12 ms cpu, 20->20->15 MB, 20 MB goal, 0 MB stacks, 0 MB globals, 10 P

一行里有五组信息:

片段含义
gc 1 @0.003s第 1 轮 GC,发生在进程启动后 0.003 秒
3%到目前为止 GC 占用的 CPU 比例
0.041+1.8+0.013 ms clock挂钟时间三段:mark setup(STW)+ 并发标记 + mark termination(STW)
0.41+0/1.8/0.036+0.13 ms cpuCPU 时间:前一个数 + assist/idle/dedicated + 后一个数
4->5->3 MB本轮开始堆大小 → 本轮结束堆大小 → 存活堆大小

关键在第四组 0.41+0/1.8/0.036+0.13:中间的 0/1.8/0.036 是辅助标记 / 后台空闲标记 / 专职标记三种标记工作的 CPU 时间。到 gc 4 变成 0.19/6.1/3.0——assist 开始出现(0.19),说明分配速率高到应用必须停下来帮 GC 干活了。

第一组 0.041+1.8+0.013 里的首尾两个数就是两次 STW:0.041 ms(mark setup)与 0.013 ms(mark termination)。中间的 1.8 ms 是并发标记,不阻塞应用。

6.2.2 实验二:调参改变的是「次数」还是「暂停」

同一负载,分别设 GOGC=100(默认)、GOGC=400、GOMEMLIMIT=64MiB,统计 2 秒内的 GC 轮数:

$ GOTOOLCHAIN=go1.27.0 GOGC=100 ./prog 2>&1 | grep -c '^gc '
24
$ GOTOOLCHAIN=go1.27.0 GOGC=400 ./prog 2>&1 | grep -c '^gc '
6
$ GOTOOLCHAIN=go1.27.0 GOMEMLIMIT=64MiB ./prog 2>&1 | grep -c '^gc '
25

GOGC 从 100 提到 400,GC 轮数从 24 降到 6(少 4 倍)——因为目标堆变成原来的 4 倍,自然要更久才触发。而 GOMEMLIMIT=64MiB 反而给出 25 轮,和默认接近:这个负载的堆远小于 64 MiB,软内存上限根本没被触到,pacing 仍按 GOGC 走。

这引出一个重要区分:GOGC 管的是「相对增长率」(目标堆 = 存活 × (1 + GOGC/100)),GOMEMLIMIT 管的是「绝对上限」。两者同时设置时取更严格的那个。但无论哪个,它们改的都是触发时机,不是单次暂停的时长——暂停时长由「标记栈大小」决定,与触发频率无关。

6.2.3 实验三:用 runtime/metrics 读暂停分布与 CPU 去向

gctrace 给的是每轮的平均值,runtime/metrics 给的是分布。程序读取几个关键指标:

import "runtime/metrics"

samples := []metrics.Sample{
	{Name: "/gc/cycles/total:gc-cycles"},
	{Name: "/gc/pauses:seconds"},
	{Name: "/cpu/classes/gc/mark/assist:cpu-seconds"},
	{Name: "/cpu/classes/gc/mark/dedicated:cpu-seconds"},
	{Name: "/cpu/classes/gc/mark/idle:cpu-seconds"},
	{Name: "/cpu/classes/gc/pause:cpu-seconds"},
	{Name: "/cpu/classes/gc/total:cpu-seconds"},
}
metrics.Read(samples)

实测输出(负载同 6.2.1,运行结束后读一次):

$ GOTOOLCHAIN=go1.27.0 ./metrics
gc cycles total        = 56
STW pause  p50         = 18µs
STW pause  p99         = 61µs
gc mark assist  (s)    = 0.0008
gc mark dedicated(s)   = 0.0212
gc mark idle     (s)   = 0.0013
gc pause         (s)   = 0.0192
gc total         (s)   = 0.0425

三点读法:

  1. 暂停的 p50 只有 18 µs,p99 也只有 61 µs。这就是「Go 的 STW 很短」的量化表达——短到对绝大多数在线服务无感。p99 是 p50 的约 3.4 倍,说明暂停有长尾,但绝对值仍在微秒级。
  2. CPU 去向里 dedicated(0.0212s)最大,assist 最小(0.0008s)。这轮负载的分配速率不算高,GC 主要由专职 worker 完成,应用几乎不用停下来辅助。如果 assist 反而变大,就是分配太快的信号(对照 6.2.1 的 gc 4)。
  3. pause(0.0192s)与 dedicated(0.0212s)同量级——注意 pause 统计的是「暂停期间所有 P 累积的 CPU 时间」,不是挂钟时间。10 个 P 各停 2 ms,累计就是 20 ms。

6.2.4 源码:阶段开关与辅助标记

阶段定义在 src/runtime/mgc.go:

const (
	_GCoff             = iota // GC not running; sweeping in background, write barrier disabled
	_GCmark                   // GC marking roots and workbufs: allocate black, write barrier ENABLED
	_GCmarktermination        // GC mark termination: allocate black, P's help GC, write barrier ENABLED
)

只有三个阶段。从 _GCoff 到 _GCmark 的切换(mark setup)和从 _GCmark 到 _GCmarktermination 的切换(mark termination)需要 STW;_GCmark 内部的标记是并发的。这就解释了 6.2.1 里 0.041+1.8+0.013 为什么首尾短、中间长——两次 STW 都很短,1.8 ms 的并发标记不阻塞。

阶段切换由 setGCPhase 完成(见 6.1.2),它在切换 gcphase 的同时开/关写屏障。阶段与屏障是绑定的:_GCmark 和 _GCmarktermination 期间屏障开启,_GCoff 期间关闭。

辅助标记的实现在 src/runtime/mgcmark.go。核心是「债务」换算——应用每分配一段内存,就欠 GC 一段扫描工作:

	// Compute the amount of scan work we need to do to make the
	// balance positive. When the required amount of work is low,
	// we over-assist to build up credit for future allocations
	// and amortize the cost of assisting.
	assistWorkPerByte := gcController.assistWorkPerByte.Load()
	assistBytesPerWork := gcController.assistBytesPerWork.Load()
	debtBytes := -gp.gcAssistBytes
	scanWork := int64(assistWorkPerByte * float64(debtBytes))
	if scanWork < gcOverAssistWork {
		scanWork = gcOverAssistWork
		debtBytes = int64(assistBytesPerWork * float64(scanWork))
	}

四个要点:

  • gp.gcAssistBytes 是每个 goroutine 的债务。分配时欠账变负,辅助标记时还账变正。
  • assistWorkPerByte 是 pacing 控制器算出的「每字节分配要还多少扫描工作」,assistBytesPerWork 是它的倒数。这两个值每轮 GC 结束时更新(见 6.2.5),不是固定常数——这是 pacing 能自适应分配速率的关键。
  • gcOverAssistWork 是「超额辅助」:当债务很小时,故意多扫一点,攒下信用(credit)抵扣未来的分配。注释写得很清楚:「amortize the cost of assisting」——把辅助成本摊销掉,避免每次小额分配都触发一次微小的辅助。
  • 还有一条「偷信用」的路径:应用可以从 gcController.bgScanCredit(后台标记攒下的信用)里偷,够的话就不用自己扫:
	bgScanCredit := gcController.bgScanCredit.Load()
	stolen := int64(0)
	if bgScanCredit > 0 {
		if bgScanCredit < scanWork {
			stolen = bgScanCredit
			gp.gcAssistBytes += 1 + int64(assistBytesPerWork*float64(stolen))
		} else {
			stolen = scanWork
			gp.gcAssistBytes += debtBytes
		}
		gcController.bgScanCredit.Add(-stolen)
		...
	}

注释说这个偷窃是「racy 的」,两个 mutator 同时偷可能把 credit 偷成负数——但「长远看无所谓」,因为负数只让后续偷窃失败,等 credit 重新累积即可。这是一处典型的「用宽松一致性换吞吐」的设计:不为了精确的信用记账付出锁成本。

标记什么时候结束?由 gcMarkDone 判断:

// gcMarkDone transitions the GC from mark to mark termination if all
// reachable objects have been marked (that is, there are no grey
// objects and can be no more in the future). Otherwise, it flushes
// all local work to the global queues where it can be discovered by
// other workers.

条件很直白:没有灰色对象、且未来也不可能有(即所有 worker 都已 drain)。因为每个 P 有自己的本地工作队列,必须先把它们「flush 到全局队列」才能确认全局也没有剩余工作——这就是为什么需要一次 STW 来做最终确认(mark termination)。

6.2.5 源码:pacing 控制器

决定「下一轮什么时候开始」的是 src/runtime/mgcpacer.go 的 gcControllerState:

type gcControllerState struct {
	// Initialized from GOGC. GOGC=off means no GC.
	gcPercent atomic.Int32

	// memoryLimit is the soft memory limit in bytes.
	// Initialized from GOMEMLIMIT. GOMEMLIMIT=off is equivalent to MaxInt64
	memoryLimit atomic.Int64

	// heapMinimum is the minimum heap size at which to trigger GC.
	// During initialization this is set to 4MB*GOGC/100.
	heapMinimum uint64

	// runway is the amount of runway in heap bytes allocated by the
	// application that we want to give the GC once it starts.
	// This is computed from consMark during mark termination.
	runway atomic.Uint64

	// consMark is the estimated per-CPU consMark ratio for the application.
	// It represents the ratio between the application's allocation
	// rate, as bytes allocated per CPU-time, and the GC's scan rate,
	// as bytes scanned per CPU-time.
	consMark float64

	// lastConsMark is the computed cons/mark value for the previous 4 GC
	// cycles.
	lastConsMark [4]float64
	...
}

逐字段读:

  • gcPercent 与 memoryLimit 就是 GOGC 与 GOMEMLIMIT 的落点。6.2.2 实验里改这两个环境变量,改的就是这两个字段。
  • heapMinimum 处理「小堆高分配」的病态场景。注释点明:当存活集很小但分配很多时,严格按 GOGC × live 触发会导致 GC 极频繁、每次的固定开销被放大。4MB*GOGC/100 的下限把这些固定开销摊薄。这是 6.2.2 里 GOGC=400 → 6 轮 之外还需要注意的一层:即使 GOGC=100,堆小于 4 MB 时也不会触发。
  • runway 是「GC 启动后应用还能分配多少」。它不是凭空的常数,而是 consMark 在 mark termination 时算出来的——pacing 的目标是让并发标记和应用分配「赛跑」时刚好同时到达终点。
  • consMark 是整套 pacing 的核心比率:(应用分配速率) / (GC 扫描速率),单位是「字节/CPU 时间」之比。注释说明它由每轮 GC 实测更新。assistWorkPerByte(6.2.4 里用的)就是从 consMark 推导出来的——这就是为什么辅助标记的量能自适应负载。
  • lastConsMark [4]float64 保留前 4 轮的值。用 4 轮的窗口做平滑,避免单轮的抖动让 pacing 剧烈摆动。

把这几个字段串起来,就是 pacing 的完整逻辑:每轮 GC 结束时实测 consMark(这轮分配了多少、扫描了多少、各花了多少 CPU),据此推算下一轮的 runway 和目标堆,再由 gcPercent/memoryLimit/heapMinimum 三个约束共同决定触发点。 应用分配越快,consMark 越大,pacing 就越早启动 GC、让辅助标记承担更多工作——这就是 6.2.1 里 gc 4 出现 assist 的机制。

6.2.6 决策:读 GC 指标的顺序

你想知道看哪个指标判读
GC 是否频繁gctrace 的行数 / /gc/cycles/total行数多但暂停短,通常是「小堆高分配」,先调 GOGC 或找分配热点
STW 有多长gctrace 首尾两数 / /gc/pauses 直方图首尾是 mark setup 与 mark termination;直方图给 p50/p99
应用有没有被拖去辅助标记gctrace 的 assist 段 / /cpu/classes/gc/mark/assistassist 显著非零 = 分配太快,pacing 让应用还债
标记 CPU 花在哪/cpu/classes/gc/mark/{assist,dedicated,idle}dedicated 是专职 worker,idle 是空闲 P 顺手做
改了 GOGC 为什么暂停没变——GOGC 只改触发时机,不改单次暂停时长
GOMEMLIMIT 设了没效果gctrace 的 goal / /gc/heap/goal:bytes堆远小于上限时上限不生效,pacing 仍按 GOGC
GC 目标堆是多少gctrace 的 N MB goal存活 × (1 + GOGC/100),受 heapMinimum 下限约束
每轮 GC 的成本趋势/gc/cycles/total 与 /cpu/classes/gc/mark/* 的变化标记 CPU 占比上升说明「分配变快 / 扫描变慢」,是 pacing 的预警(consMark 本身未对外导出)

三条读指标纪律:

  1. 先分「频率」与「单次成本」。GOGC/GOMEMLIMIT 调的是频率;单次成本由标记栈大小与分配速率决定。混淆这两者会导致「调了参数没效果」的困惑(6.2.2 的实测就是例子)。
  2. STW 看直方图,不看平均值。p50 好看不代表没有长尾,/gc/pauses 的 p99 才是线上服务的真实体验。
  3. assist 非零要当信号,不是当错误。它意味着 pacing 判断「应用分配得比 GC 扫得快」,是设计内的行为;要改的是分配速率,不是去关掉 assist。

下一节把 6.2 的观察方法用在一件具体的事上:Green Tea GC 在 1.26 与 1.27 上的实测差异。

阅读导航:上一节:6.1 三色标记与写屏障 · 下一节:6.3 Green Tea GC 实测对比 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「golang」更多文章

  1. 《Go 语言编程实战》目录
  2. 《Go 语言编程实战》18.3 上线、观测与迭代
  3. 《Go 语言编程实战》18.2 故障演练