本节目标:搞清 Actuator 端点的暴露与鉴权策略,理解 Micrometer 的
MeterRegistry模型,能写出可用的业务指标与公共标签,并让 Prometheus 真正算出 P99。
适用版本:Spring Boot 4.1.x(Java 21)
16.1 Actuator 与指标
15.3 节给 book-loan 配好了健康检查与优雅停机,服务能上线了。但「能启动」和「能运营」是两回事:线上流量一上来,你需要的是一组能回答「现在健康吗、慢不慢、借书量多少」的数字。这正是 Actuator 加 Micrometer 要解决的事。
本节先讲端点该怎么开(这是安全问题,不是配置问题),再讲指标模型与自定义指标,最后落到 Prometheus 抓取与开箱指标的读法。16.2 讲链路追踪,16.3 讲日志聚合,三节合起来才是完整的可观测性。
端点暴露:默认只有 health
引入 spring-boot-starter-actuator 后,Boot 并不会把一堆端点敞开。Web 端点的默认暴露集合只有一个:
management:
endpoints:
web:
exposure:
include: health # 这是默认值,不用写也是它
这个默认值来自自动配置元数据里的 management.endpoints.web.exposure.include(默认 ['health'])。也就是说,/actuator/beans、/actuator/env、/actuator/heapdump 这些端点虽然被注册了,但默认不通过 HTTP 暴露。
要让 Prometheus 抓到指标,必须显式放行 prometheus;要临时调日志级别,得放行 loggers。生产上的原则是只暴露必要端点,且一律要求鉴权。下面这张表列出几个「看着无害、实则高危」的端点:
| 端点 | 返回内容 | 风险 |
|---|---|---|
/actuator/env | 全部 PropertySource 与属性值 | 泄漏数据库地址、账号,脱敏不彻底时连密码一起给 |
/actuator/configprops | 所有 @ConfigurationProperties | 同上,且结构更完整 |
/actuator/heapdump | 一份完整堆转储(HPROF) | 含内存里的明文 token、密钥;且大堆 dump 会停顿 JVM |
/actuator/threaddump | 全部线程栈 | 暴露内部类名、SQL、调用链 |
/actuator/beans | 全部 bean 及依赖关系 | 泄漏内部结构 |
/actuator/loggers(POST) | 可改日志级别 | 被改成 TRACE 会打爆磁盘与 CPU |
/actuator/shutdown | 关闭应用 | 默认禁用,一旦开启就是远程关机 |
env 与 configprops 默认对敏感值做脱敏(show-values 默认 never),但脱敏是基于键名匹配的,业务自定义的密钥字段经常不在规则里。不要把「它脱敏了」当成安全保证。
推荐的生产配置是「单独管理端口 + 最小暴露集 + 独立鉴权」:
management:
server:
port: 8081 # 管理端点走单独端口,不进业务网关
endpoints:
web:
exposure:
include: health,info,prometheus,metrics,loggers
endpoint:
health:
show-details: when-authorized
loggers:
enabled: true
管理端口 8081 通常只对集群内网开放,业务流量走 8080。对 loggers 这种可写端点,必须叠加鉴权——复用 9.1 Security 配置模型
里的 SecurityFilterChain,给 /actuator/** 单开一条链要求 ADMIN 角色即可。别把 actuator 全放开:这是历史上最常见的生产事故来源之一。
/actuator/metrics 与 MeterRegistry 模型
Micrometer 的核心抽象是 io.micrometer.core.instrument.MeterRegistry。所有指标都是注册到它上面的 Meter。Boot 一旦在 classpath 上发现 Micrometer 与某个 registry 实现,就会自动装配一个 MeterRegistry bean,并把它注入到需要的地方。
Micrometer 只有四种基本计量器,理解它们的语义比记 API 重要:
| 计量器 | 接口 | 语义 | 典型用途 | Prometheus 侧 |
|---|---|---|---|---|
| 计数器 | Counter | 只增不减的累计值 | 请求数、借出次数、错误数 | _total |
| 仪表 | Gauge | 瞬时采样值,可增可减 | 当前在借数、连接池活跃数、队列长度 | 直接 gauge |
| 计时器 | Timer | 时长分布 + 次数 + 总时长 | 借书耗时、HTTP 延迟、DB 查询 | _count / _sum / _bucket |
| 分布摘要 | DistributionSummary | 非时间量的分布 | 请求体大小、借阅本数 | 同 Timer,单位非秒 |
Counter 与 Gauge 的区别是「累计」与「瞬时」;Timer 与 DistributionSummary 的区别是「时间量」与「非时间量」。选错类型会让告警逻辑变得别扭——比如用 Gauge 记请求数,重启就归零,速率算不出来。
/actuator/metrics 列出当前所有指标名,/actuator/metrics/{name} 查看某个指标的可用标签与当前值。下面是 book-loan 暴露后的示例输出(不同实例的数值不同,这里只示意结构):
$ curl -s localhost:8081/actuator/metrics | python3 -m json.tool
{
"names": [
"hikaricp.connections.active",
"http.server.requests",
"jvm.memory.used",
"bookloan.loans.created",
"bookloan.loan.duration"
]
}
$ curl -s "localhost:8081/actuator/metrics/bookloan.loans.created" | python3 -m json.tool
{
"name": "bookloan.loans.created",
"measurements": [ { "statistic": "COUNT", "value": 1287.0 } ],
"availableTags": [
{ "tag": "branch", "values": ["main", "east"] },
{ "tag": "application", "values": ["book-loan"] }
]
}
注意 availableTags:Micrometer 会列出该指标上出现过的标签值。标签值是基数的来源——每个不同的标签组合都会生成一条独立的时间序列,这一点在后面讲 MeterFilter 时会成为重点。
自定义业务指标
框架开箱指标只能告诉你「服务慢不慢」,回答不了「业务做得好不好」。book-loan 至少要两个业务指标:借出次数(Counter)与借书耗时(Timer)。
不要在每个类里各自 new 一个计数器,统一注入 MeterRegistry,用工厂方法按名字加标签取计量器。Micrometer 对同名同标签的计量器是幂等的,重复取不会重复注册:
package com.example.loan.metrics;
import io.micrometer.core.instrument.Counter;
import io.micrometer.core.instrument.MeterRegistry;
import io.micrometer.core.instrument.Timer;
import java.util.concurrent.TimeUnit;
import org.springframework.stereotype.Component;
@Component
public class LoanMetrics {
private final MeterRegistry registry;
public LoanMetrics(MeterRegistry registry) {
this.registry = registry;
}
/** 借出次数:只增不减,用 Counter */
public void recordLoanCreated(String branch) {
Counter.builder("bookloan.loans.created")
.description("成功创建的借阅单数量")
.tag("branch", branch)
.register(registry)
.increment();
}
/** 借书耗时:带时长分布,用 Timer */
public Timer.Sample startTimer() {
return Timer.start(registry);
}
public void stopTimer(Timer.Sample sample, String result) {
sample.stop(Timer.builder("bookloan.loan.duration")
.description("借书请求处理耗时")
.tag("result", result) // success / conflict / rejected
.publishPercentileHistogram() // 生成 Prometheus 直方图桶
.register(registry));
}
}
Timer.Sample 的用法值得单独说:它记录开始时刻,在结束时才把时长写进指标,中间那段业务代码不需要持有 Timer 对象。这对「一次请求里跨多个方法测量」很方便。等价地,也可以用 timer.record(() -> doBorrow()) 包住一段同步代码,或 recordCallable 处理有返回值和异常的场景。
再看一个 Gauge 的例子——「当前在借数量」是瞬时值,必须用 Gauge。Gauge 的关键点是它持有一个对象的引用并每次抓取时回调取值,而不是推一个数字进去:
// 用一个可变的持有者承载「当前在借数」,Gauge 每次抓取时回调读取
public void bindActiveLoans(java.util.concurrent.atomic.AtomicInteger active) {
io.micrometer.core.instrument.Gauge.builder("bookloan.loans.active", active, java.util.concurrent.atomic.AtomicInteger::get)
.description("当前处于借出状态的图书数量")
.tag("branch", "main")
.register(registry);
}
常见错误是把 Gauge 当成「可以 set 的值」:直接 registry.gauge("x", 5) 注册一个常量,之后无论业务怎么变它永远是 5。Gauge 必须绑定到一个会被更新的对象(AtomicInteger、集合的 size() 方法引用等),否则它只会输出注册那一刻的快照。
命名上建议点分隔 + 单位后缀:累计量用 .created / .total,时长用 .duration,大小用 .bytes,数量用 .active / .pending。Micrometer 会给时间类自动换算单位,但名字里带单位后缀(如 .duration)能让 Prometheus 查询更直观。
MeterFilter 与公共标签
线上服务通常有多个实例、多套环境,指标必须能按「应用」和「实例」切开,否则聚合在一起毫无意义。两个公共标签是标配:
management:
metrics:
tags:
application: book-loan
application 标签由 management.metrics.tags.* 统一加到所有指标上。instance 标签则不要写死在配置文件里——它必须每个实例不同,通常由部署时注入(Kubernetes 里用 POD_NAME,或启动参数 -Dmanagement.metrics.tags.instance=${HOSTNAME})。写死成固定值会导致多实例指标互相覆盖。
比公共标签更重要的,是控制基数。这是 Micrometer 使用中最大的坑:标签值的每一种组合都是一条独立时间序列,如果把 userId、loanId 这种高基数维度做成标签,指标数会爆炸,Prometheus 内存被吃光。防御手段是注册 MeterFilter bean:
package com.example.loan.metrics;
import io.micrometer.core.instrument.config.MeterFilter;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
@Configuration
public class MetricsConfig {
@Bean
MeterFilter commonTagsFilter() {
return MeterFilter.commonTags(io.micrometer.core.instrument.Tags.of(
"application", "book-loan"));
}
@Bean
MeterFilter cardinalityGuard() {
// 单指标标签组合上限,超过后新组合被丢弃,防止基数爆炸拖垮 registry
return MeterFilter.maximumAllowableTags("bookloan.loans.created", "branch", 100,
MeterFilter.deny());
}
@Bean
MeterFilter dropNoisy() {
// 拒绝掉不需要的高频内部指标,减少抓取体积
return MeterFilter.denyNameStartsWith("jvm.buffer.");
}
}
MeterFilter 的常用工厂方法:commonTags(...) 加公共标签、deny() / accept() 按条件丢弃或放行、denyNameStartsWith(...)、maximumAllowableMetrics(int) 限制总指标数、maximumAllowableTags(...) 限制单指标标签组合数。它们按 bean 注册顺序链式生效。
一条经验规则:标签值应该是「有限的、可枚举的」——状态码、结果类型、机构编号、HTTP 方法都可以;用户 ID、订单号、请求参数不行。需要下钻到单笔请求时,那是追踪(16.2)该干的事,不是指标。
让 Prometheus 能算 P99:percentiles-histogram
这是很多人第一次接 Prometheus 时踩的坑:配好了抓取,却发现 P99 算不出来,只有 _count 和 _sum。
原因是 Micrometer 的 Timer 默认只发布计数与总和,不发布直方图桶。而 Prometheus 的 histogram_quantile() 函数必须基于 _bucket 序列才能算分位数。要让 Micrometer 发布桶,必须显式打开:
management:
metrics:
distribution:
percentiles-histogram:
http.server.requests: true # 只给 HTTP 指标开
bookloan.loan.duration: true # 业务耗时也开
management.metrics.distribution.percentiles-histogram 是一个 Map<String, Boolean>,键是指标名(可带前缀匹配),值是否开启。打开后 Micrometer 会生成一组固定的桶边界,Prometheus 端就能算:
histogram_quantile(0.99,
sum(rate(bookloan_loan_duration_seconds_bucket[5m])) by (le)
)
另一种方式是在客户端直接算分位数(management.metrics.distribution.percentiles),让应用算出 P50/P90/P99 后作为 gauge 上报。两者不要混用,但要知道它们的取舍:
| 方式 | 计算位置 | 能否跨实例聚合 | 序列数 | 适用 |
|---|---|---|---|---|
percentiles-histogram | Prometheus 服务端 | 能,by (le) 后聚合 | 较高(每个桶一条) | 需要全局 P99、跨实例告警 |
percentiles | 应用客户端 | 不能,只能各自算 | 低 | 单实例内部分析、非精确场景 |
客户端分位数无法跨实例聚合——把两个实例各自的 P99 求平均,得到的不是全局 P99。所以线上告警应该走 percentiles-histogram,让 Prometheus 在服务端聚合。
/actuator/prometheus 端点与抓取
Prometheus 是拉模型,需要应用暴露一个能被抓取的文本端点。这需要两样东西:micrometer-registry-prometheus 依赖,以及放行 prometheus 端点。
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
<dependency>
<groupId>io.micrometer</groupId>
<artifactId>micrometer-registry-prometheus</artifactId>
</dependency>
依赖到位后,/actuator/prometheus 会返回 Prometheus 文本格式的指标。下面是对应的抓取配置(prometheus.yml):
scrape_configs:
- job_name: book-loan
metrics_path: /actuator/prometheus
static_configs:
- targets: ['book-loan-0:8081', 'book-loan-1:8081']
scrape_interval: 15s
抓取端点的示例输出(节选,格式为 Prometheus exposition 文本):
# HELP bookloan_loans_created_total 成功创建的借阅单数量
# TYPE bookloan_loans_created_total counter
bookloan_loans_created_total{branch="main",application="book-loan",} 1287.0
bookloan_loans_created_total{branch="east",application="book-loan",} 402.0
# HELP bookloan_loan_duration_seconds 借书请求处理耗时
# TYPE bookloan_loan_duration_seconds histogram
bookloan_loan_duration_seconds_bucket{result="success",le="0.05",} 1203.0
bookloan_loan_duration_seconds_bucket{result="success",le="0.1",} 1270.0
bookloan_loan_duration_seconds_bucket{result="success",le="+Inf",} 1287.0
bookloan_loan_duration_seconds_count{result="success",} 1287.0
bookloan_loan_duration_seconds_sum{result="success",} 42.13
命名规则:Micrometer 的 bookloan.loans.created 被翻译成 bookloan_loans_created_total(点转下划线、Counter 加 _total 后缀);Timer 生成 _bucket / _count / _sum 三条。抓取配置里 metrics_path 与暴露端口必须对上,否则 Prometheus 报 404。
开箱指标怎么读
Boot 自动装配了三类最有价值的开箱指标,读懂它们能覆盖大部分日常排障:
| 指标 | 含义 | 该看什么 |
|---|---|---|
jvm.memory.used{area="heap"} | 堆内存使用 | 是否持续爬升(泄漏)还是锯齿(正常 GC) |
jvm.gc.pause | GC 停顿分布 | P99 停顿是否超过 SLA;Full GC 是否频繁 |
jvm.threads.live | 存活线程数 | 是否持续增长(线程池泄漏) |
http.server.requests | 请求延迟与状态码 | 按 uri/status/outcome 切分看 P99 与错误率 |
hikaricp.connections.active | 池内活跃连接 | 是否逼近 maximum-pool-size |
hikaricp.connections.pending | 等待连接的线程数 | 长期大于 0 说明池太小或连接泄漏 |
hikaricp.connections.timeout | 获取连接超时次数 | 大于 0 即需告警 |
HikariCP 的指标是自动绑定的:只要 classpath 上有 HikariCP 与 MeterRegistry,Boot 的 DataSourcePoolMetricsAutoConfiguration 就会挂上 HikariDataSourceMeterBinder,无需手写代码。http.server.requests 的标签由 management.metrics.web.server.max-uri-tags 控制(默认 100),超出后会归并到 uri="UNKNOWN"——自定义 URL 带路径参数时要注意基数,否则很快打满。连接池参数的调优见 11.1 HikariCP 调优
。
常见坑
- 把 actuator 全暴露:
include: "*"等于把env、heapdump挂到公网,是真实数据泄漏。 Gauge注册成常量:值永远停在注册那一刻,业务变了指标不动。- 用高基数维度做标签:
userId、loanId会让时间序列数爆炸,Prometheus 内存被打爆。 instance标签写死:多实例指标互相覆盖,聚合结果错误。- 没开
percentiles-histogram却想算 P99:只有_count/_sum,histogram_quantile无桶可用。 - 客户端
percentiles拿来跨实例告警:分位数不能平均,全局 P99 必须在 Prometheus 端算。 loggers端点无鉴权:被改成 TRACE 会打爆磁盘,必须叠加角色控制。
小结
Actuator 的默认暴露集只有 health,生产上应「单独管理端口 + 最小暴露集 + 独立鉴权」,env 与 heapdump 尤其危险。指标的核心是 MeterRegistry 上的四类计量器:Counter、Gauge、Timer、DistributionSummary,选型看语义是累计还是瞬时、是时间量还是非时间量。自定义指标通过工厂方法按名字加标签幂等注册,公共标签与 MeterFilter 负责可聚合性与基数防御。要让 Prometheus 算 P99,必须打开 percentiles-histogram 生成桶,且在服务端聚合。JVM、HTTP、HikariCP 的开箱指标开箱即用,是日常排障的第一站。
阅读导航:上一节:15.3 健康检查与优雅停机 · 下一节:16.2 链路追踪 。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。