4.2 逃逸分析与分配决策
go build -gcflags='-m' 打印的每一行,都是一次编译期推断的结论。大多数人对它的用法停留在「看到 escapes to heap 就去改代码」——但如果不知道这行是谁写的、依据是什么,就很容易改错方向:把本该逃逸的对象硬塞回栈,或者为了「不逃逸」写出更难维护的代码。
本节不重复「什么写法会逃逸」的清单(那是卷三 6.1 的内容),而是从编译器的数据流分析实现入手,回答「这一行是谁打印的、它凭什么这么判断」。
本节要回答:
-m的输出每一行对应编译器里哪一步、-m -m多出来的flow:轨迹是什么?结论是:逃逸分析在src/cmd/compile/internal/escape/里先把函数体建成「位置—洞(location/hole)」数据流图,再用leaks位图记录每个位置到堆/到返回值的最短解引用距离,最后在solve.go里迭代到不动点;-m打印结论,-m -m额外打印explainFlow生成的路径。 与卷三《Go 语言高级编程》6.1 的分工:卷三写的是应用侧决策表(什么写法会逃逸、怎么改代码),本节写的是编译器实现(数据流图与位图编码),只写增量。同样,/posts/golang/下讲逃逸的专题文章偏概念,本节只补「输出逐行怎么读」。
4.2.1 实验:三种 -m 的输出怎么读
复现基线:
- Go 工具链
go version go1.27.0 darwin/arm64(GOTOOLCHAIN=go1.27.0) - 机器:Apple M1 Pro,10 核,32 GiB
- 分析对象是纯编译期行为,不受
GOGC/GOMAXPROCS影响,也不需要运行程序 - 命令均为
go build -gcflags=... -o /dev/null .;-m是单级,-m -m是两级,-l关闭内联
被测代码刻意覆盖四类典型情形:返回结构体值、返回局部变量地址、装箱到 any、闭包捕获变量。
type Point struct{ X, Y int }
func stackAlloc() Point {
p := Point{1, 2}
return p
}
func heapAlloc() *Point {
p := Point{3, 4}
return &p
}
func leakToInterface() any {
v := 42
return v
}
func closureCapture() func() int {
x := 0
return func() int { x++; return x }
}
第一遍,默认内联,单级 -m。注意输出里混着内联决策,这是最容易被误读的地方:
$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m' -o /dev/null .
# escape
./main.go:7:6: can inline stackAlloc
./main.go:12:6: can inline heapAlloc
./main.go:17:6: can inline leakToInterface
./main.go:22:6: can inline closureCapture
./main.go:24:9: can inline closureCapture.func1
./main.go:27:6: can inline sliceLiteral
./main.go:32:24: inlining call to stackAlloc
./main.go:32:38: inlining call to heapAlloc
...
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap
./main.go:23:2: moved to heap: x
./main.go:24:9: func literal escapes to heap
./main.go:28:14: []int{...} escapes to heap
./main.go:32:13: ... argument does not escape
./main.go:32:24: ~r0 escapes to heap
./main.go:32:28: *(~r0) escapes to heap
can inline / inlining call to 来自内联阶段,moved to heap / escapes to heap / does not escape 来自逃逸分析。两者共享同一个 -m 开关,这是第一件要知道的事。
第二遍,加 -l 关闭内联,让逃逸结论「裸奔」——这样输出里只剩逃逸分析自己的判断:
$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m -l' -o /dev/null .
# escape
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap
./main.go:23:2: moved to heap: x
./main.go:24:9: func literal escapes to heap
./main.go:28:14: []int{...} escapes to heap
./main.go:32:13: ... argument does not escape
./main.go:32:24: stackAlloc() escapes to heap
./main.go:32:28: *heapAlloc() escapes to heap
./main.go:32:77: .autotmp_0() escapes to heap
./main.go:32:93: sliceLiteral(1) escapes to heap
对比两遍输出可以看到一个关键事实:同一段 main,开内联时打印的是 ~r0 escapes to heap(内联后 stackAlloc 的返回值没有名字,用临时 ~r0 表示),关内联时打印的是 stackAlloc() escapes to heap。内联会把「跨函数的逃逸」变成「函数内的逃逸」,因此 -m 的输出会随内联决策而变。想看清函数内部的原始判断,必须先 -l。
第三遍,两级 -m -m,多出 flow: 轨迹。这是本节最有用的一档,它把「为什么」摊开:
$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m -m' -o /dev/null .
./main.go:31:6: cannot inline main: function too complex: cost 218 exceeds budget 80
./main.go:13:2: p escapes to heap in heapAlloc:
./main.go:13:2: flow: ~r0 ← &p:
./main.go:13:2: from &p (address-of) at ./main.go:14:9
./main.go:13:2: from return &p (return) at ./main.go:14:2
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap in leakToInterface:
./main.go:19:9: flow: ~r0 ← &{storage for 42}:
./main.go:19:9: from 42 (spill) at ./main.go:19:9
./main.go:19:9: from return 42 (return) at ./main.go:19:2
./main.go:23:2: closureCapture capturing by ref: x (addr=false assign=true width=8)
./main.go:23:2: x escapes to heap in closureCapture:
./main.go:23:2: flow: {storage for func literal} ← &x:
./main.go:23:2: from x (captured by a closure) at ./main.go:24:22
./main.go:23:2: from x (reference) at ./main.go:24:22
flow: 一行是路径的终点(~r0 ← &p 读作「返回值 ~r0 收到了 p 的地址」),下面缩进的 from ... (原因) at 位置 是逐跳的边。整段合起来就是一条从「取地址」到「泄漏到返回值」的证据链。(address-of)、(spill)、(captured by a closure)、(slice-literal-element) 这些括号里的词是编译器给每条边打的标签。
4.2.2 源码:数据流图、leaks 位图、不动点求解
逃逸分析的入口在 src/cmd/compile/internal/escape/escape.go。Funcs 自底向上遍历函数,Batch 对一批函数做分析:
func Funcs(all []*ir.Func) {
reassignOracles := make(map[*ir.Func]*ir.ReassignOracle)
ir.VisitFuncsBottomUp(all, func(list []*ir.Func, recursive bool) {
Batch(list, reassignOracles)
})
}
Batch 的注释直接点明了它做三件事:建图、流闭包、解不动点:
func Batch(fns []*ir.Func, reassignOracles map[*ir.Func]*ir.ReassignOracle) {
var b batch
b.heapLoc.attrs = attrEscapes | attrPersists | attrMutates | attrCalls
...
// Construct data-flow graph from syntax trees.
for _, fn := range fns {
b.initFunc(fn)
}
for _, fn := range fns {
if !fn.IsClosure() {
b.walkFunc(fn)
}
}
...
b.walkAll()
b.finish(fns)
}
图的基本单元在 src/cmd/compile/internal/escape/graph.go:location 是「一个可能有地址的位置」(变量、临时值、返回值……),hole 是「位置上的一个洞」(带解引用偏移和备注),edge 连接两者。hole 上的 addr/deref 操作会移动偏移:
func (k hole) shift(delta int) hole {
n := k
n.derefs += delta
...
return n
}
func (k hole) deref(where ir.Node, why string) hole { return k.shift(1).note(where, why) }
func (k hole) addr(where ir.Node, why string) hole { return k.shift(-1).note(where, why) }
why 就是 4.2.1 输出里括号里的 (address-of)、(spill) 这些标签——它们是边在建立时被记下来的,不是打印时猜的。
每个位置的「逃逸程度」用 src/cmd/compile/internal/escape/leaks.go 的 leaks 位图编码。它是一个 [8]uint8,下标含义固定:
const (
leakHeap = iota
leakMutator
leakCallee
leakResult0
)
// Heap returns the minimum deref count of any assignment flow from l
// to the heap. If no such flows exist, Heap returns -1.
func (l leaks) Heap() int { return l.get(leakHeap) }
注意 Heap() 的语义是**「到堆的最短解引用次数」**,不是布尔值。get 用「减一」编码,所以「没有路径」是 -1:
func (l leaks) get(i int) int { return int(l[i]) - 1 }
func (l *leaks) add(i int, derefs int) {
if old := l.get(i); old < 0 || derefs < old {
l.set(i, derefs)
}
}
add 只在「新的路径更短」时更新——这就是为什么结论是「最短距离」。Optimize 再砍掉比堆路径更长的结果路径:
func (l *leaks) Optimize() {
if x := l.Heap(); x >= 0 {
for i := 1; i < len(*l); i++ {
if l.get(i) >= x {
l.set(i, -1)
}
}
}
}
求解在 src/cmd/compile/internal/escape/solve.go:walkAll。它把 allLocs、heapLoc、mutatorLoc、calleeLoc 全部推入队列,反复传播直到不动点:
func (b *batch) walkAll() {
todo := newQueue(len(b.allLocs) + 3)
...
for _, loc := range b.allLocs {
todo.pushFront(loc)
loc.queuedWalkAll = true
}
todo.pushFront(&b.mutatorLoc)
todo.pushFront(&b.calleeLoc)
todo.pushFront(&b.heapLoc)
...
for todo.len() > 0 {
root := todo.popFront()
root.queuedWalkAll = false
walkgen++
b.walkOne(root, walkgen, enqueue)
}
}
-m -m 的 flow: 行就来自这里。solve.go:explainFlow 沿 src.dst 一路回溯到根,按 base.Flag.LowerM >= 2 决定是否打印:
func (b *batch) explainFlow(pos string, dst, srcloc *location, derefs int, notes *note, explanation []*logopt.LoggedOpt) []*logopt.LoggedOpt {
ops := "&"
if derefs >= 0 {
ops = strings.Repeat("*", derefs)
}
print := base.Flag.LowerM >= 2
flow := fmt.Sprintf(" flow: %s ← %s%v:", b.explainLoc(dst), ops, b.explainLoc(srcloc))
if print {
fmt.Printf("%s:%s\n", pos, flow)
}
...
}
ops 的构造逻辑说明了一件事:flow: 里的 & 或 * 数量就是那条边的解引用增量。~r0 ← &p 里只有一个 &,代表「取一层地址」。
最终打印「逃逸/不逃逸」结论的是 escape.go:reportLeaks:
func (b *batch) reportLeaks(pos src.XPos, name string, esc leaks, sig *types.Type) {
warned := false
if x := esc.Heap(); x >= 0 {
if x == 0 {
base.WarnfAt(pos, "leaking param: %v", name)
} else {
base.WarnfAt(pos, "leaking param content: %v", name)
}
warned = true
}
...
if !warned {
base.WarnfAt(pos, "%v does not escape", name)
}
}
esc.Heap() == 0 打印 leaking param,> 0 打印 leaking param content——这两个词的差别就是解引用层数:前者是把指针本身漏出去,后者是把指针指向的内容漏出去。
最后,还有一类逃逸与数据流无关,由 HeapAllocReason 单独判定(src/cmd/compile/internal/escape/utils.go),例如「过大的局部数组」「make 出来的切片长度非常量」。这类在 Batch 里被直接接到 heapLoc:
if why := HeapAllocReason(loc.n); why != "" {
b.flow(b.heapHole().addr(loc.n, why), loc)
}
4.2.3 决策:从 -m 输出反推改法
-m 输出里的字样 | 编译器里的来源 | 该怎么读 / 怎么改 |
|---|---|---|
moved to heap: p | reportLeaks + esc.Heap(),变量本身要堆化 | 该变量的地址被存进了堆对象或返回;想留栈上就得断开这条 flow: 链 |
42 escapes to heap | 常量/临时值被装箱(leaks 的 leakHeap) | 典型是装进 any 或闭包;改法是避免接口装箱(如预分配、泛型) |
... argument does not escape | 变参切片未逃逸 | 说明 fmt 这类调用没把变参留在堆上;不用改 |
~r0 escapes to heap | 内联后的匿名返回值 | 开内联才出现的形态;用 -l 看原始判断 |
flow: ~r0 ← &p | explainFlow 打印的路径 | 终点在左、起点在右;顺着缩进的 from 找到「谁取了这个地址」 |
leaking param: x | esc.Heap() == 0 | 指针本身漏出去;改函数签名(返回副本而非指针) |
leaking param content: x | esc.Heap() > 0 | 指针指向的内容漏出去;通常无害,别急着改 |
x does not escape | reportLeaks 的兜底分支 | 理想状态;这是结论,不是建议 |
cannot inline main: cost 218 exceeds budget 80 | 内联成本模型 | 与逃逸无直接关系,但内联会改变逃逸结论;分析前先看有没有 cannot inline |
closureCapture capturing by ref: x | flowClosure(escape.go) | 闭包按引用捕获;by ref 意味着 x 会逃逸 |
三条实操建议:
- 分析前先加
-l。内联会把跨函数逃逸折叠成函数内逃逸,输出形态完全不同;先用-m -l看原始判断,再用-m -m看路径。 - 优先看
flow:,而不是结论行。结论行只告诉你「逃逸了」,flow:告诉你「从哪一步漏出去的」——后者才是能改的地方。 - 别为了消灭
escapes to heap而牺牲可读性。leaking param content这类「内容逃逸」在很多场景下是良性的;把它当成 bug 去修,往往是把清晰代码改成晦涩代码。真正的收益点是热路径上的高频分配——先用 4.3 的 profile 找到热点,再回来用本节的方法逐行读。
阅读导航:上一节:4.1 size class 与 mcache/mcentral/mheap · 下一节:4.3 分配热点定位与对象复用 。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。