Erlang 是动态类型语言,但这并不意味着类型只能靠运行时校验。Dialyzer(DIscrepancy AnalYZer for ERlang programs)通过静态分析在编译期发现类型不一致、不可达代码、不可能成功的模式匹配等缺陷,是 Erlang 生态中最有价值的质量工具之一。它的独特之处在于不做类型推断的强制约束,而是基于 success typing 计算「函数在什么输入下可能成功」,因此既不会拒绝合法程序,也不会放过明显的类型矛盾。本文将系统讲解 Dialyzer 的原理、typespec 语法、PLT 管理、告警治理以及与测试的协同。
一、Dialyzer 与 success typing 原理
1.1 与传统类型检查的区别
传统静态类型系统(如 Haskell、Rust)采用「悲观」策略:无法证明类型正确就拒绝编译。Dialyzer 采用「乐观」策略:只报告那些能证明一定出错的代码。
| 维度 | 传统类型检查 | Dialyzer success typing |
|---|---|---|
| 判定标准 | 能否证明类型安全 | 能否证明一定失败 |
| 对合法程序的干扰 | 可能误报 | 几乎无误报 |
| 对缺陷的覆盖 | 完整 | 仅覆盖可证明的部分 |
| 是否需要标注 | 强制 | 可选,无标注也能分析 |
1.2 success typing 的核心思想
Dialyzer 为每个函数计算一个「成功类型」——即如果该函数返回了值,那么输入与输出必然属于某个集合。例如:
%% 没有 spec 也能分析
add(X, Y) -> X + Y.
%% Dialyzer 推导出的 success typing 大致为:
%% add(number(), number()) -> number()
如果调用处传入 add("a", 1),Dialyzer 会发现 + 运算符作用于字符串与整数时必然抛 badarith,于是报告:
Function add/2 has no local return
「no local return」意味着该调用不可能正常返回,这正是 success typing 的表达方式。
1.3 分析的能力边界
Dialyzer 能发现的问题:
- 类型矛盾(把
atom()传给期望integer()的位置); - 不可达代码(
case分支永远匹配不到); - 不可能成功的模式匹配;
- 函数无本地返回(必然崩溃);
- 未定义的函数调用(配合 xref)。
Dialyzer 不能发现的问题:
- 逻辑错误(算错了但类型正确);
- 并发竞态;
- 消息协议不匹配(除非显式标注类型);
- 未处理的边界值(只要类型上合法)。
二、类型规范语法:spec 与 type
2.1 函数 spec
-module(payment).
%% 单子句 spec
-spec charge(Amount :: pos_integer(), Currency :: atom()) -> {ok, binary()} | {error, term()}.
%% 多子句用分号分隔
-spec parse(binary()) -> {ok, map()} | {error, invalid_format}.
%% 带约束的 spec
-spec split(non_neg_integer(), [T]) -> {[T], [T]} when T :: term().
2.2 自定义类型
-module(types).
%% 类型别名
-type user_id() :: pos_integer().
-type email() :: binary().
-type status() :: active | inactive | suspended.
%% 记录类型
-record(user, {
id :: user_id(),
email :: email(),
tags :: [atom()]
}).
-type user() :: #user{}.
%% 不透明类型:外部只能整体使用,不能拆解
-opaque handle() :: {ref, reference(), node()}.
%% 导出类型供其他模块引用
-export_type([user_id/0, user/0, handle/0]).
| 关键字 | 用途 | 可见性 |
|---|---|---|
-type | 公开类型别名 | 需 -export_type 才能跨模块 |
-opaque | 不透明类型 | 外部无法依赖内部结构 |
-spec | 函数签名 | 随模块导出 |
-callback | behaviour 回调签名 | 由 behaviour 模块声明 |
2.3 内置类型速查
%% 基础
atom(), binary(), bitstring(), boolean(), float(), integer(),
number(), pid(), port(), reference(), tuple(), map(), list(),
nil(), term(), any(), none(), no_return()
%% 参数化
[Type] % 元素为 Type 的列表
[Type, ...] % 非空列表
#{key := Type} % 必含 key 的 map
#{key => Type} % 可含 key 的 map
fun((A) -> B) % 函数类型
fun((...) -> B) % 任意参数
tuple() | {A, B} % 定长元组
%% 数值区间
1..255
0..16#FF
2.4 用 spec 描述协议
对消息传递型系统,spec 是唯一能约束消息格式的手段:
-type request() :: {get, binary()} | {put, binary(), term()}.
-type reply() :: {ok, term()} | {error, term()}.
-spec handle_call(request(), gen_server:from(), state()) ->
{reply, reply(), state()}.
实践:把跨进程消息定义为
-type,并让handle_call/3、handle_cast/2、handle_info/2的 spec 引用它。Dialyzer 会据此检查发送方与接收方是否一致。
三、PLT 构建与增量分析
3.1 什么是 PLT
PLT(Persistent Lookup Table)是 Dialyzer 预计算的类型信息库,包含 OTP 标准库、第三方依赖与自身代码的 success typing。第一次运行需要为每个模块构建 PLT,耗时可能达数分钟。
%% rebar.config
{dialyzer, [
{warnings, [unmatched_returns, error_handling, underspecs]},
{plt_apps, all_deps}, % PLT 包含所有依赖
{plt_extra_apps, [cowboy, jsx]},
{plt_location, local},
{plt_prefix, "my_service"}
]}.
3.2 构建与更新
%% 首次构建(含所有依赖,较慢)
rebar3 dialyzer
%% 依赖变更后重建 PLT
rebar3 dialyzer --update-plt
%% 只分析不重建
rebar3 dialyzer --no_check_plt
3.3 PLT 缓存策略
| 场景 | 做法 | 效果 |
|---|---|---|
| 本地开发 | PLT 放 _build/default/ | 增量更新 |
| CI | 缓存 _build/*/rebar3_*_plt | 省去重建 |
| 多分支 | 按 OTP 版本分 key | 避免 PLT 冲突 |
| Docker | PLT 预置进镜像 | 构建提速数分钟 |
# GitHub Actions 缓存 PLT
- uses: actions/cache@v4
with:
path: _build/default/rebar3_*_plt
key: plt-${{ runner.os }}-${{ hashFiles('rebar.lock') }}
3.4 增量分析
Dialyzer 支持基于模块的增量分析:只重新分析变更的模块及其调用者。这依赖 PLT 中保存的依赖图,因此不要在每次构建时删除 PLT,否则每次都退化为全量分析。
四、常见告警解读与治理
4.1 no local return
Function validate/1 has no local return
含义:该函数在任何输入下都不会正常返回,必然抛异常或死循环。常见原因是内部调用了必然失败的操作:
%% 错误示范:maps:get 无默认值 + 后续解构必然失败
validate(Config) ->
Timeout = maps:get(timeout, Config), % 若 Config 是 list 则必然 badmap
Timeout * 1000.
%% 修复:加类型守卫或默认值
validate(Config) when is_map(Config) ->
Timeout = maps:get(timeout, Config, 5000),
Timeout * 1000.
4.2 类型矛盾
The call payment:charge(_amount::binary(), _currency::atom())
breaks the contract (Amount :: pos_integer(), Currency :: atom())
原因:调用方传入的类型与 spec 声明不符。治理方式有二——修调用方,或修正 spec:
%% 若确实需要支持字符串金额,改 spec 并加转换
-spec charge(Amount :: pos_integer() | binary(), Currency :: atom()) ->
{ok, binary()} | {error, term()}.
charge(Amount, Currency) when is_binary(Amount) ->
charge(binary_to_integer(Amount), Currency);
charge(Amount, Currency) when is_integer(Amount), Amount > 0 ->
do_charge(Amount, Currency).
4.3 unmatched returns
开启 unmatched_returns 后,Dialyzer 会报告返回值被忽略的调用,常用于发现忘记检查的错误:
%% 告警:返回值被忽略
{ok, _} = application:ensure_all_started(my_app), % OK,已匹配
file:write_file("/tmp/x", Data), % 告警:返回值未处理
%% 修复
case file:write_file("/tmp/x", Data) of
ok -> continue();
{error, Reason} -> logger:error("write failed: ~p", [Reason])
end.
4.4 underspecs 与 overspecs
| 告警 | 含义 | 处理 |
|---|---|---|
underspecs | spec 比实际行为更宽泛 | 收窄 spec |
overspecs | spec 比实际行为更严格 | 放宽 spec 或补分支 |
specdiffs | spec 与推导类型有差异 | 复核后修正 |
%% underspecs:实际只返回 ok | error,spec 写得太宽
-spec save(map()) -> term(). % 太宽
-spec save(map()) -> ok | {error, term()}. % 准确
4.5 告警白名单与抑制
对确实无法修复(如第三方库缺陷)的告警,用 -dialyzer 属性局部抑制:
%% 抑制单个函数的特定告警
-dialyzer({nowarn_function, legacy_adapter/2}).
%% 抑制特定类型的告警
-dialyzer({nowarn_function, [old_api/1, old_api/2]}).
%% 抑制整模块的某类告警(慎用)
-dialyzer(no_return).
原则:抑制必须带注释说明原因与到期时间,否则告警会随时间累积成「技术债黑洞」。
五、与测试及 CI 的配合
5.1 静态分析与测试的分工
| 缺陷类型 | 首选工具 |
|---|---|
| 类型矛盾、不可达代码 | Dialyzer |
| 未定义函数、废弃调用 | xref |
| 业务逻辑错误 | EUnit / Common Test |
| 并发与属性 | PropEr / QuickCheck |
| 代码风格 | elvis / rebar3_lint |
5.2 与属性测试互补
属性测试(见 https://plumephp.com/elixir-testing-property/)负责「用随机输入找反例」,Dialyzer 负责「用静态推导证明不可能」。两者覆盖不同维度:
%% 属性测试发现的具体反例,往往能反哺 spec 修正
prop_reverse() ->
?FORALL(L, list(integer()),
lists:reverse(lists:reverse(L)) =:= L).
%% 对应的 spec 应精确表达
-spec reverse([T]) -> [T] when T :: term().
5.3 CI 集成与门禁
- name: Static Analysis
run: |
rebar3 xref
rebar3 dialyzer
%% rebar.config —— 让告警成为构建失败条件
{dialyzer, [
{warnings, [unmatched_returns, error_handling, underspecs,
unknown, race_conditions]},
{plt_apps, all_deps}
]}.
落地节奏:遗留项目一次性开启全部告警会产生海量输出。建议先用
nowarn_function抑制存量,把门禁加在新代码上,然后按月收敛存量。
5.4 从 dialyzer 输出定位根因
%% 生成详细报告
rebar3 dialyzer > dialyzer.log 2>&1
%% 按模块聚合告警数量
grep -oP '(?<=^)\w+(?=\.erl)' dialyzer.log | sort | uniq -c | sort -rn
把告警数量纳入质量看板,能直观看到技术债的收敛趋势。
六、总结
Dialyzer 是 Erlang 工程化中最被低估的工具。它的 success typing 模型决定了它「宁可漏报、绝不误报」,因此可以放心地在 CI 中作为硬门禁。落地要点:
- 理解模型:success typing 只报告可证明的失败,不追求完备性;
- 写 spec:用
-spec/-type/-opaque表达意图,让 Dialyzer 有据可依; - 管 PLT:PLT 必须缓存,绝不在每次构建时重建,否则耗时不可接受;
- 治告警:先分类(no local return / 类型矛盾 / unmatched returns),再逐个修复,抑制必须带注释;
- 进 CI:与 xref、测试、属性测试组成互补的质量矩阵。
把 Dialyzer 用好之后,Erlang 项目的类型缺陷会在编译期而非生产事故中暴露,这与 https://plumephp.com/erlang-production-cases/ 中强调的「把故障左移」理念一脉相承。下一步可以结合 https://plumephp.com/erlang-behaviour-custom/ 中的回调契约,让 behaviour 的 -callback 也被 Dialyzer 校验。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。