《Go 语言运行时原理》2.1 G/M/P 结构与状态机

G/M/P 是什么,站内 Go 专题已经讲过。本节只写增量:用 GODEBUG=scheddetail=1 的真实回显,把每个 goroutine 的 status 与 waitreason 逐字段读出来,再定位到 src/runtime/runtime2.go 的状态常量与 g/m/p 结构体,最后给出一张可直接对照的状态迁移表。

从这一章开始进入调度器。先立规矩:GMP 的三个字母分别代表什么、g/m/p 有哪些字段,站内 Go 专题(/posts/golang/)里的介绍型文章已经讲得很充分,本节不复述结构定义。本节只写增量——用 scheddetail 的真实回显把「一个 goroutine 此刻处于什么状态、为什么」逐字段读出来,再把这些状态对回 runtime2.go 里的常量与结构体。

本节要回答:scheddetail 输出里每个 status= 和括号里的词,对应源码里的哪个常量、哪次状态迁移? 结论是:scheddetail 里的 status=0/1/2/3/4/6 就是 _Gidle/_Grunnable/_Grunning/_Gsyscall/_Gwaiting/_Gdead;括号里的文字来自 waitReason 表;掌握了这两张表的对应,就能从一行回显直接判断「这个 goroutine 卡在哪」。

2.1.1 实验:用 scheddetail 看真实状态

准备一个制造多种状态的程序:24 个 CPU 密集型 goroutine,外加运行时的后台 goroutine(GC、sysmon 等):

package main

import (
	"runtime"
	"sync"
	"time"
)

//go:noinline
func burn(n int) uint64 {
	var x uint64 = 1
	for i := 0; i < n; i++ {
		x = x*1664525 + 1013904223
	}
	return x
}

func main() {
	var wg sync.WaitGroup
	var sum [24]uint64
	for w := 0; w < 24; w++ {
		wg.Add(1)
		go func(id int) {
			defer wg.Done()
			deadline := time.Now().Add(2500 * time.Millisecond)
			for time.Now().Before(deadline) {
				sum[id] = burn(50000)
			}
		}(w)
	}
	wg.Wait()
	runtime.KeepAlive(sum)
}

程序刚启动、还没开始忙时,scheddetail 长这样(截取 SCHED 0ms 这一段):

GODEBUG=schedtrace=1500,scheddetail=1 ./sched
SCHED 0ms: gomaxprocs=10 idleprocs=7 threads=4 spinningthreads=1 needspinning=0 idlethreads=1 runqueue=0 gcwaiting=false nmidlelocked=-1 stopwait=0 sysmonwait=false
  P8: status=1 schedtick=0 syscalltick=0 m=3 runqsize=0 gfreecnt=0 timerslen=0
  P9: status=0 schedtick=2 syscalltick=0 m=nil runqsize=0 gfreecnt=0 timerslen=0
  M3: p=8 curg=nil mallocing=0 throwing=0 preemptoff= locks=17 dying=0 spinning=false blocked=false lockedg=nil
  M2: p=nil curg=nil mallocing=0 throwing=0 preemptoff= locks=0 dying=0 spinning=false blocked=true lockedg=nil
  M1: p=nil curg=nil mallocing=0 throwing=0 preemptoff= locks=32 dying=0 spinning=false blocked=false lockedg=nil
  M0: p=nil curg=nil mallocing=0 throwing=0 preemptoff= locks=17 dying=0 spinning=false blocked=false lockedg=1
  G1: status=1(chan receive) m=nil lockedm=0
  G2: status=4(force gc (idle)) m=nil lockedm=nil
  G3: status=4(GC sweep wait) m=nil lockedm=nil
  G4: status=1() m=nil lockedm=nil

把这一段读出来,能得到四条信息:

  • P8: status=1 ... m=3 表示 P8 正被 M3 占用(_Prunning),而 P9: status=0 ... m=nil 表示 P9 空闲(_Pidle)。
  • G1: status=1(chan receive) 是 _Grunnable + waitreason=chan receive——它刚从 channel 上被唤醒、还没排上 CPU。
  • G2: status=4(force gc (idle))、G3: status=4(GC sweep wait) 是 _Gwaiting,即运行时后台 goroutine 在等 GC 相关工作。
  • G4: status=1() 是 _Grunnable 且无等待原因,就是普通的就绪用户 goroutine。

程序跑起来、10 个 P 全忙之后,同一程序的状态分布变成(截取 SCHED 1511ms):

SCHED 1511ms: gomaxprocs=10 idleprocs=0 threads=11 spinningthreads=0 needspinning=1 idlethreads=0 runqueue=10 gcwaiting=false nmidlelocked=0 stopwait=0 sysmonwait=false
  P0: status=1 schedtick=47 syscalltick=0 m=4 runqsize=0 gfreecnt=0 timerslen=0
  P1: status=1 schedtick=46 syscalltick=0 m=10 runqsize=1 gfreecnt=0 timerslen=0
  G1: status=4(sync.WaitGroup.Wait) m=nil lockedm=nil
  G2: status=4(force gc (idle)) m=nil lockedm=nil
  G17: status=4(GOMAXPROCS updater (idle)) m=nil lockedm=nil
  G18: status=1() m=nil lockedm=nil
  G23: status=2() m=8 lockedm=nil
  G28: status=2() m=10 lockedm=nil
  G34: status=2() m=6 lockedm=nil
  G37: status=2() m=9 lockedm=nil

对比可以看出一条清晰的规律:status=2(_Grunning)的 goroutine 一定有 m=<非 nil>;status=1(_Grunnable)和 status=4(_Gwaiting)的 m 都是 nil。 这正是状态定义要求的——只有 _Grunning 才持有 M 和栈的所有权。G1: status=4(sync.WaitGroup.Wait) 则是 main 协程卡在 wg.Wait() 上。

复现基线:Go 1.27.0 darwin/arm64;Apple M1 Pro,10 逻辑核,32 GiB 内存;GOMAXPROCS=10(默认),GOGC=100。程序参数:24 个 goroutine,每个循环 burn(50000),持续 2.5 秒;采样周期 1500ms,共采到 2 个样本点。

2.1.2 源码:状态常量与三个结构体

scheddetail 里那个数字,是 g.atomicstatus 的取值。它的常量定义在 src/runtime/runtime2.go:

// src/runtime/runtime2.go:31(gstatus 常量片段)
const (
	_Gidle = iota // 0
	_Grunnable // 1
	_Grunning  // 2
	_Gsyscall  // 3
	_Gwaiting  // 4
	_Gmoribund_unused // 5
	_Gdead // 6
	_Genqueue_unused // 7
	_Gcopystack // 8
)

括号里的文字来自 waitReason 表(同文件 :1221 起)。几个最常见的取值:

// src/runtime/runtime2.go:1224(waitReason 片段)
waitReasonZero                  waitReason = iota // ""
waitReasonGCAssistWait                            // "GC assist wait"
waitReasonSelect                                  // "select"
waitReasonChanReceive                             // "chan receive"
waitReasonSyncWaitGroupWait                       // "sync.WaitGroup.Wait"

所以 status=4(chan receive) 就是 atomicstatus==_Gwaiting 且 waitreason==waitReasonChanReceive。运行时在 proc.go 的 schedtrace 里把这两者拼起来打印(proc.go:7029),这也是为什么 status=1() 的括号是空的——_Grunnable 没有等待原因。

三个结构体都在 runtime2.go:type g struct(:471)、type m struct(:616)、type p struct(:774),全局调度器状态 type schedt struct(:932)。状态迁移则集中在 proc.go 的几个函数里:

grep -n "^func gopark\|^func goready\|^func execute\|^func goschedImpl\|^func goexit0" proc.go
457:func gopark(unlockf func(*g, unsafe.Pointer) bool, lock unsafe.Pointer, reason waitReason, traceReason traceBlockReason, traceskip int) {
493:func goready(gp *g, traceskip int) {
3346:func execute(gp *g, inheritTime bool) {
4322:func goschedImpl(gp *g, preempted bool) {
4506:func goexit0(gp *g) {

迁移本身由 casgstatus 完成,它是个原子 CAS。proc.go 里能数出这几条主干迁移:

3360:	casgstatus(gp, _Grunnable, _Grunning)   // execute:即将运行
4290:	casgstatus(gp, _Grunning, _Gwaiting)    // gopark:主动阻塞
1145:	casgstatus(gp, _Gwaiting, _Grunnable)   // goready:被唤醒

execute 把 _Grunnable 推到 _Grunning,gopark 把 _Grunning 退回 _Gwaiting,goready 再把 _Gwaiting 推回 _Grunnable。这三条构成了 goroutine 生命周期的骨架;其余几十处 casgstatus 都是这三条在具体场景(channel、锁、GC、syscall)里的特化。

g 结构体里与状态直接相关的字段有四个,读调度代码时盯住它们就够:

// src/runtime/runtime2.go:471(g 结构体片段)
type g struct {
	stack       stack   // 该 goroutine 的栈区间 [lo, hi)
	m           *m      // 当前绑定的 M(仅 _Grunning/_Gsyscall 非空)
	atomicstatus atomic.Uint32 // 上面那张 gstatus 表的取值
	goid         uint64        // goroutine id,scheddetail 里的 "G23" 就是它
	waitsince    int64      // 进入阻塞的近似时刻
	waitreason   waitReason // 阻塞原因,scheddetail 括号里的文字
	preempt      bool       // 抢占信号,重复 stackguard0 = stackpreempt
}

scheddetail 的 G23: status=2() m=8 逐个字段对下来就是:goid=23、atomicstatus=_Grunning、waitreason=waitReasonZero、m.id=8。

P 也有自己的状态机,取值在 runtime2.go:132 起:_Pidle(0)、_Prunning(1)、_Psyscall_unused(2,已废弃)、_Pgcstop(3)、_Pdead(4)。所以 scheddetail 里 P8: status=1 是 _Prunning、P9: status=0 是 _Pidle。P 的状态只有持它的 M 能改,_Pgcstop 是 STW 时所有 P 的统一归宿。

全局的 schedt 结构体(runtime2.go:932)则装着那些「全局唯一」的计数:goidgen(goroutine id 发生器)、nmidle/nmidlelocked(空闲 M 数)、nmspinning(正在自旋找活的 M 数)、runq(全局运行队列)、gcwaiting(GC 是否在等 STW)。scheddetail 第一行的 spinningthreads=/needspinning=/runqueue= 就是从这里读的。

2.1.3 决策:状态迁移表与排障对照

把实验和源码合成两张表。第一张是状态迁移表,用来回答「一个 goroutine 是怎么动起来的」:

起始状态触发目标状态迁移函数(文件:函数)
_Gidle (0)新 goroutine 初始化_Grunnable (1)proc.go:newproc
_Grunnable (1)被调度执行_Grunning (2)proc.go:execute
_Grunning (2)阻塞(channel/锁/GC)_Gwaiting (4)proc.go:gopark
_Grunning (2)被抢占/主动让出_Grunnable (1)proc.go:goschedImpl
_Grunning (2)进入系统调用_Gsyscall (3)proc.go:entersyscall
_Gwaiting (4)被唤醒_Grunnable (1)proc.go:goready
_Grunning (2)函数返回、退出_Gdead (6)proc.go:goexit0
_Grunning (2)栈增长需复制_Gcopystack (8)stack.go:copystack

第二张表把常见的 waitReason 翻译成「发生了什么、该看哪一节」:

scheddetail 里的词含义排查方向
chan receive / chan send卡在 channel 收发见 3.2 的 trace 时间线
sync.WaitGroup.Wait卡在 wg.Wait()检查是否有 goroutine 未 Done
select卡在 select检查所有 case 是否都不满足
force gc (idle)GC 强制触发后台协程正常,GC 相关
GC sweep wait / GC scavenge waitGC 清扫/回收后台协程正常,GC 相关
GOMAXPROCS updater (idle)自动调整 GOMAXPROCS 的后台协程见 3.1、3.3

三条使用纪律:

  1. 先看 status=2 的 goroutine 有几个,是否等于 idleprocs 的补数。 忙碌状态下 status=2 的数量应当接近 gomaxprocs;远小于说明大量 goroutine 在 _Gwaiting。
  2. status=1 堆积(runqueue 变大)说明 CPU 是瓶颈。 如 2.1.1 里 runqueue=10、idleprocs=0,就是典型的「活儿多、核少」。
  3. status=4 的括号是排障的第一线索。 它直接告诉你阻塞在哪个原语上,省去翻代码的功夫。

还有两条从结构体推出来的判断规则,值得记住:

  • _Grunning 的 goroutine 数不可能超过 gomaxprocs。 因为只有拿到 P 的 M 才能把 goroutine 置为 _Grunning,而 P 的总数是 gomaxprocs。如果你在 scheddetail 里看到 status=2 的行数大于 gomaxprocs,那一定是采样时状态在并发变化,不是真有那么多在跑。
  • m=nil 的 _Grunnable goroutine 一定在某个运行队列里。 要么在某个 P 的本地 runq(runqsize),要么在全局 sched.runq(第一行的 runqueue=)。两者相加应约等于所有 status=1 的 goroutine 数——这是校验采样是否自洽的一个小技巧。

最后提醒一个采样特性:schedtrace 的每一次采样都要 lock(&sched.lock) 并逐个读取所有 P/M/G,它本身会干扰被测程序,采样越密干扰越大。所以它适合观察趋势,不适合精确计时。

P 的状态取值也整理成一张对照表,方便和 scheddetail 的 P0: status=? 对齐:

P 状态数值含义出现时机
_Pidle0空闲,未绑定 M无活可干、等待唤醒
_Prunning1被 M 占用、正在跑用户代码正常忙碌
_Psyscall_unused2已废弃(历史遗留)不再出现
_Pgcstop3为 STW 停下GC 的 STW 阶段
_Pdead4不再使用(GOMAXPROCS 调小)动态调整 GOMAXPROCS 后

在 2.1.1 的忙碌样本里,10 个 P 全是 status=1(_Prunning),idleprocs=0;空闲样本里 idleprocs=7,对应的就是 7 个 _Pidle 的 P。

一句话收束本节:G 是任务,M 是执行任务的线程,P 是执行任务所需的资源(本地队列 + 缓存);状态机描述的是 G 在「就绪—运行—阻塞」之间的流转,而 M 和 P 只是承载这个流转的容器。 下一节看容器怎么把任务高效地分发出去。

记住状态常量比记住结构体字段更重要:0/1/2/3/4/6 这六个数字会在你调试的每一个 scheddetail、每一份 trace、每一次 panic 栈里反复出现。

阅读导航:上一节:1.3 复现基线:环境与基准约定 · 下一节:2.2 调度循环与 work stealing 实测 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「golang」更多文章

  1. 《Go 语言编程实战》目录
  2. 《Go 语言编程实战》18.3 上线、观测与迭代
  3. 《Go 语言编程实战》18.2 故障演练