全文检索的质量在写入那一刻就已经决定了大半:文档被切成哪些词、是否做了归一化、同义词有没有展开,直接决定了之后能召回什么。Elasticsearch 把这一过程抽象成分析器,由字符过滤器、分词器、token 过滤器三段串成一条流水线。内置分析器能覆盖多数场景,但中文分词、拼音检索、业务同义词这些需求必须自己组装。本文从三段式结构讲起,覆盖各段的可选组件、拼音与同义词配置、analyze 调试,以及索引期与查询期分析器分离的设计原则。
1. 分析器的三段式结构
一句话总结: 分析器由字符过滤器、分词器、token 过滤器依次串联,前一段的输出是后一段的输入。
1.1 三段的职责
- 字符过滤器(character filter):在切词之前处理原始字符串,做整体替换,例如去掉 HTML 标签、把
&换成and。 - 分词器(tokenizer):把字符串切成一个个 token,同时记录每个 token 的起止偏移量。
- token 过滤器(token filter):对 token 流做逐个处理,可以小写化、去停用词、加同义词、做词干还原。
三段中只有分词器是必需的,字符过滤器与 token 过滤器都可以没有。一个分析器最多一个分词器、可以零到多个字符过滤器与 token 过滤器。
1.2 内置分析器的对照
# 用 analyze API 观察不同内置分析器的切词差异
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{
"analyzer": "standard",
"text": "The Quick Brown-Fox 2026"
}'
standard 会输出 the、quick、brown、fox、2026,连字符被当分隔符。若换成 keyword,整句作为一个 token。若换成 simple,数字 2026 会被丢掉。理解这些差异是选型的前提。
1.3 索引期与查询期的分离
analyzer 决定文档怎么切,search_analyzer 决定查询词怎么切。默认查询沿用索引分析器,但在需要「索引细、查询粗」或反之的场景,必须显式分开设置,后文会专门展开。
2. 字符过滤器
一句话总结: 字符过滤器在切词前对原文做替换,常见的是 HTML 剥离与字符映射。
2.1 html_strip
curl -X PUT "localhost:9200/articles" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"analyzer": {
"html_analyzer": {
"type": "custom",
"char_filter": ["html_strip"],
"tokenizer": "standard",
"filter": ["lowercase", "stop"]
}
}
}
},
"mappings": {
"properties": {
"content": { "type": "text", "analyzer": "html_analyzer" }
}
}
}'
html_strip 会把 <b>粗体</b> 处理成 粗体,同时把 & 解码成 &。
2.2 mapping 字符过滤器
curl -X PUT "localhost:9200/products" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"char_filter": {
"normalize_symbols": {
"type": "mapping",
"mappings": ["٠ => 0", "١ => 1", "٢ => 2", "- => -", "+ => +"]
}
},
"analyzer": {
"product_analyzer": {
"type": "custom",
"char_filter": ["normalize_symbols"],
"tokenizer": "standard",
"filter": ["lowercase"]
}
}
}
}
}'
mapping 过滤器支持 => 单字符映射,也支持用逗号分隔的等价组(如 "a,b => c" 表示 a 与 b 都映射为 c)。
2.3 pattern_replace 正则替换
{
"char_filter": {
"remove_dashes": {
"type": "pattern_replace",
"pattern": "(\\d)-(\\d)",
"replacement": "$1$2"
}
}
}
上面的规则把 138-0013-8000 归一化成 13800138000,适合电话号码检索。正则必须小心灾难性回溯,避免在大文本上卡死。
3. 分词器的选型
一句话总结: 中文必须用专门分词器,英文默认 standard 已够用,结构化字段用 keyword。
3.1 standard 与 whitespace 的取舍
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "tokenizer": "whitespace", "text": "hello,world foo-bar" }'
whitespace 输出 hello,world 与 foo-bar 两个 token,标点原样保留。这适合日志、代码片段这类「标点是语义一部分」的场景。
3.2 中文分词器
# 安装 IK 插件后定义自定义词典路径
curl -X PUT "localhost:9200/news" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"analyzer": {
"ik_smart_custom": { "type": "custom", "tokenizer": "ik_smart",
"filter": ["lowercase"] },
"ik_max_word_custom": { "type": "custom", "tokenizer": "ik_max_word",
"filter": ["lowercase"] }
}
}
},
"mappings": {
"properties": {
"title": { "type": "text", "analyzer": "ik_max_word",
"search_analyzer": "ik_smart" }
}
}
}'
IK 提供两种模式:ik_max_word 做最细粒度切分,召回高但索引大;ik_smart 做粗粒度切分,精度高但召回略低。经典组合是索引用 max_word、查询用 smart,兼顾召回与精度。
3.3 keyword 与 pattern 分词器
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{
"tokenizer": { "type": "pattern", "pattern": "[/\\\\-]" },
"text": "2026/10/01-release"
}'
输出 2026、10、01、release。pattern 分词器适合路径、版本号这类有固定分隔符的字段。
4. Token 过滤器链
一句话总结: token 过滤器是分析器里最灵活的一段,小写、停用词、词干、ngram 都在这里完成。
4.1 常用过滤器
curl -X PUT "localhost:9200/docs" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"filter": {
"english_stop": { "type": "stop",
"stopwords": ["the", "a", "an", "of"] },
"english_stem": { "type": "stemmer", "language": "english" }
},
"analyzer": {
"english_custom": { "type": "custom", "tokenizer": "standard",
"filter": ["lowercase", "asciifolding", "english_stop", "english_stem"] }
}
}
}
}'
过滤器顺序影响结果:先小写再去停用词,才能匹配到小写形式的停用词表。asciifolding 把 café 转成 cafe,对多语言混排很有用。
4.2 ngram 与 edge_ngram
curl -X PUT "localhost:9200/autocomplete" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"filter": {
"edge_ngram_filter": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 }
},
"analyzer": {
"autocomplete_index": { "type": "custom", "tokenizer": "standard",
"filter": ["lowercase", "edge_ngram_filter"] },
"autocomplete_search": { "type": "custom", "tokenizer": "standard",
"filter": ["lowercase"] }
}
}
},
"mappings": {
"properties": {
"name": { "type": "text", "analyzer": "autocomplete_index",
"search_analyzer": "autocomplete_search" }
}
}
}'
这是索引期与查询期分离的经典案例:索引期把 elasticsearch 切成 el、ela、elas……,查询期只用普通分词,于是输入 ela 就能命中。
4.3 过滤器链的性能代价
ngram 的 max_gram 每加 1,索引体积近似线性增长。同义词展开会显著放大 token 数量。生产上要在召回效果与索引成本之间实测取平衡。
5. 同义词与拼音
一句话总结: 同义词提升召回,拼音支持首字母与全拼检索,两者都通过 token 过滤器挂到分析器上。
5.1 同义词过滤器
curl -X PUT "localhost:9200/synonyms_demo" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"filter": {
"my_synonyms": { "type": "synonym",
"synonyms": ["手机, 移动电话, 蜂窝电话", "笔记本 => 笔记本电脑",
"es, elasticsearch"] }
},
"analyzer": {
"synonym_analyzer": { "type": "custom", "tokenizer": "ik_smart",
"filter": ["lowercase", "my_synonyms"] }
}
}
}
}'
a, b, c 是等价语法,三者互相映射;a => b 是单向映射,只在查询时把 a 换成 b。生产上词典通常放在 synonyms.txt 里,用 synonyms_path 引用,并配合 _reload_search_analyzers 热更新:
curl -X POST "localhost:9200/synonyms_demo/_reload_search_analyzers?pretty"
5.2 拼音检索
curl -X PUT "localhost:9200/contacts" -H 'Content-Type: application/json' -d '
{
"settings": {
"analysis": {
"filter": {
"pinyin_filter": { "type": "pinyin", "keep_first_letter": true,
"keep_full_pinyin": true, "keep_original": true,
"remove_duplicated_term": true }
},
"analyzer": {
"pinyin_analyzer": { "type": "custom", "tokenizer": "keyword",
"filter": ["pinyin_filter"] }
}
}
},
"mappings": {
"properties": {
"name": { "type": "text", "analyzer": "ik_smart",
"fields": { "pinyin": { "type": "text", "analyzer": "pinyin_analyzer" } } }
}
}
}'
这里用了 multi-field:name 走中文分词,name.pinyin 走拼音分析器。查询时用 multi_match 同时打两个字段,就能同时支持中文与拼音输入。
5.3 同义词与拼音的组合顺序
顺序错会导致同义词失效:同义词词典里写的是汉字,若先转拼音,汉字形态已经不存在了。因此 filter 数组里 synonym 必须在 pinyin 之前。
6. analyze API 调试
一句话总结: analyze API 是分析器调试的唯一权威工具,能看到每一步的 token 与偏移量。
6.1 用指定分析器观察
curl -X POST "localhost:9200/synonyms_demo/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "analyzer": "synonym_analyzer", "text": "手机" }'
返回的每个 token 都带 position、start_offset、end_offset,可以用来核对同义词是否真的展开了。
6.2 逐段验证
# 只验证分词器
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "tokenizer": "ik_smart", "text": "中华人民共和国" }'
# 验证过滤器链(不给分词器则默认 standard)
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "filter": ["lowercase", "my_synonyms"], "text": "ES 与 Elasticsearch" }'
6.3 用 explain 模式看中间结果
curl -X POST "localhost:9200/synonyms_demo/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "field": "name", "text": "手机", "explain": true }'
6.4 用 termvectors 看已索引的词
curl -X GET "localhost:9200/contacts/_termvectors/1?fields=name.pinyin&pretty"
当怀疑「analyze 结果对但搜不到」时,termvectors 往往能揭示真相:文档里实际索引的 token 与预期不符。
7. 索引期与查询期分析器分离
一句话总结: 索引期追求召回,查询期追求精度,两者分离是自定义分析器设计的核心原则。
7.1 为什么必须分离
以自动补全为例:写入 elasticsearch 要切成 el、ela 等前缀,查询 ela 却不应被切成前缀,否则会匹配到所有以 el 开头的词。若两者用同一个分析器,效果必然打折。
7.2 配置方式
{
"properties": {
"title": { "type": "text", "analyzer": "ik_max_word",
"search_analyzer": "ik_smart" },
"name": { "type": "text", "analyzer": "autocomplete_index",
"search_analyzer": "autocomplete_search" }
}
}
7.3 多字段策略
curl -X PUT "localhost:9200/multi_field_demo" -H 'Content-Type: application/json' -d '
{
"mappings": {
"properties": {
"content": { "type": "text", "analyzer": "ik_smart",
"fields": {
"keyword": { "type": "keyword", "ignore_above": 256 },
"pinyin": { "type": "text", "analyzer": "pinyin_analyzer" }
} }
}
}
}'
一个字段同时支持中文分词检索、精确匹配与拼音检索,代价是索引体积变大。按实际查询模式裁剪子字段,是控制成本的关键。
8. 总结
| 环节 | 要点 |
|---|---|
| 三段式结构 | 字符过滤器整体替换、分词器切词、token 过滤器逐个改写 |
| 字符过滤器 | html_strip 剥离标签,mapping 做符号归一,pattern_replace 正则替换 |
| 分词器选型 | 中文用 IK 或 jieba,结构化字段用 keyword,路径用 pattern |
| 过滤器链 | 小写、asciifolding、停用词、词干、ngram 依次生效,顺序敏感 |
| 同义词 | 等价与映射两种语法,外置词典可热更新 |
| 拼音 | pinyin 插件转全拼与首字母,常与中文分析器做成 multi-field |
| 调试手段 | analyze 看切词,termvectors 看实际索引,explain 看逐段变化 |
| 索引查询分离 | 索引期保召回、查询期保精度,用 search_analyzer 与多字段实现 |
分析器是检索质量的源头,选错分词器或漏配同义词,后面再多的打分调优也补不回来。把三段流水线搭对、用 analyze API 验证、用多字段兼顾多种查询模式,文本检索就成功了一半。下一篇我们讲如何用搜索模板把查询参数化,让复杂 DSL 可以复用与集中管理。
延伸阅读
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。