「数据建模与 Mapping 设计」

讲解 ES 数据建模:Mapping 字段类型与 dynamic 策略、analyzer 多字段、nested 与父子对比、Doc Values 聚合,以及字段不可变与 reindex 重构。

ES 不是 NoSQL 数据库的简单替代,它把同一份数据拆成倒排索引、Doc Values、BKD 树等多种结构,Mapping 一旦创建就很难改。因此数据建模必须前置思考:字段用 text 还是 keyword,嵌套怎么建模,聚合怎么存,这些都直接决定检索能力与性能上限。

1. Mapping 基础

1.1 Mapping 是什么

Mapping 是索引的字段结构定义,等价于关系库的表结构,描述每个字段的类型、分析器、是否索引、是否存储 Doc Values。

概念ES 对应关系库类比
索引Index数据库/表
文档Document行
字段Field列
MappingMapping表结构定义
分析器Analyzer无(全文索引)

1.2 显式创建 Mapping

curl -X PUT 'http://localhost:9200/products?pretty' \
  -H 'Content-Type: application/json' \
  -d '{
    "mappings": {
      "properties": {
        "name":    { "type": "text", "analyzer": "ik_max_word" },
        "brand":   { "type": "keyword" },
        "price":   { "type": "float" },
        "stock":   { "type": "integer" },
        "on_sale": { "type": "boolean" },
        "tags":    { "type": "keyword" },
        "created_at": { "type": "date", "format": "yyyy-MM-dd HH:mm:ss" }
      }
    }
  }'

1.3 字段类型总览

分类类型用途
全文text分词检索
精确keyword过滤/排序/聚合
数值long/integer/float/half_float数值范围
时间date日期范围/排序
结构object/nested嵌套对象
特殊geo_point/ip/join地理/网络/父子

2. 核心字段类型

2.1 text 与 keyword 的选择

对比textkeyword
索引方式分词整体单值
匹配match 全文term 精确
聚合不支持支持
排序不支持支持
典型值标题、正文状态、品牌、ID
{
  "properties": {
    "status": { "type": "keyword" },
    "title":  { "type": "text", "analyzer": "ik_max_word" }
  }
}

最常见的建模错误是把所有字段设成 text,导致排序聚合全不可用,或把所有字段设成 keyword 导致无法全文检索。规则是:需要分词匹配的用 text,需要精确过滤与聚合的用 keyword。

2.2 数值与浮点精度

ES 数值类型遵循 Lucene BKD 树,但求和与统计依赖 Doc Values。金额类数据建议用 scaled_float 或存分为单位的 integer,避免 float 精度误差。

{
  "properties": {
    "amount_cents": { "type": "integer" },
    "rating":       { "type": "half_float" }
  }
}

2.3 date 与 format

{
  "properties": {
    "created_at": { "type": "date", "format": "strict_date_optional_time||epoch_millis" }
  }
}

date 字段底层存为自 epoch 的毫秒数,format 只影响读写展示。范围查询、排序、日期直方图聚合都依赖正确的 date 类型,字符串存 date 是常见反模式。

3. dynamic 策略与显式映射

3.1 三种 dynamic 行为

策略行为适用
true自动推断并添加字段开发期
false忽略新字段,不索引字段可扩展但不用
strict遇到新字段直接拒绝严格 schema
{
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "name": { "type": "text" }
    }
  }
}

3.2 自动推断的坑

dynamic 为 true 时,ES 会为第一个文档推断类型,同名字段后续类型不一致会被拒绝或产生索引错误。

{
  "mappings": {
    "dynamic_templates": [
      {
        "strings_as_keyword": {
          "match_mapping_type": "string",
          "mapping": { "type": "keyword" }
        }
      },
      {
        "longs_as_long": {
          "match_mapping_type": "long",
          "mapping": { "type": "long" }
        }
      }
    ]
  }
}

生产环境建议 dynamic 设为 false 或 strict,用 dynamic_templates 兜底常见字段形态,避免字段类型漂移。

3.3 字段命名规范

字段名使用 snake_case 或 camelCase 统一风格,避免字段名过长。ES 不限制字段数量,但字段越多 mapping 与查询开销越大,动辄上千字段的文档应评估是否扁平化。

4. Analyzer 与多字段设计

4.1 为每个字段配分析器

{
  "properties": {
    "title": {
      "type": "text",
      "fields": {
        "keyword": { "type": "keyword", "ignore_above": 256 },
        "ik":      { "type": "text", "analyzer": "ik_max_word" },
        "en":      { "type": "text", "analyzer": "english" }
      }
    }
  }
}

multi-fields 让同一字段以多种方式索引:title 默认 standard 分词,title.keyword 支持精确排序,title.ik 支持中文切分,title.en 支持英文词干。查询时按需选择子字段。

4.2 中文场景的多字段

{
  "mappings": {
    "properties": {
      "content": {
        "type": "text",
        "analyzer": "ik_max_word",
        "search_analyzer": "ik_smart"
      }
    }
  }
}

索引侧用 ik_max_word 全切分保召回,查询侧用 ik_smart 精切分保精度,是中文搜索的标准做法。分析器配置方法详见《倒排索引与分词原理》。

4.3 ignore_above 与空值

keyword 字段设置 ignore_above 后,超过长度的值不索引,避免超长字符串撑爆词典。空字符串、null 数组在聚合与过滤中的行为不同,建模时要明确默认值语义。

5. 嵌套与父子文档

5.1 object 的扁平化陷阱

{
  "properties": {
    "reviews": {
      "type": "nested",
      "properties": {
        "author": { "type": "keyword" },
        "rating": { "type": "integer" },
        "content": { "type": "text" }
      }
    }
  }
}

默认 object 类型把嵌套对象扁平化为独立字段,数组对象会丢失对象边界,例如两个评价的 author 与 rating 会交叉匹配。需要保持对象独立性时用 nested。

5.2 nested 与 join 对比

维度nestedjoin(父子)
数据存储同文档隐藏块父子独立文档
更新粒度整篇重写子文档独立更新
查询能力nested_queryhas_child/has_parent
性能快,单分片内较慢,跨文档
父子数量无限制父文档扇出有限
{
  "mappings": {
    "properties": {
      "question": { "type": "join", "relations": { "question": "answer" } }
    }
  }
}

5.3 建模选型建议

一对多且不频繁更新的内嵌列表用 nested;子文档大量独立更新(如订单明细状态)用 join。join 查询 has_child 性能开销大,千级扇出即可能劣化,能扁平化就扁平化。

6. 数值与聚合建模

6.1 Doc Values 与聚合

{
  "properties": {
    "price": {
      "type": "float",
      "doc_values": true
    }
  }
}

聚合与排序依赖 Doc Values 列存储。默认开启,若确认某字段只做检索不做聚合排序,可关闭 doc_values 省磁盘;反之需聚合的字段必须开启。

6.2 稀疏字段与成本

文档间字段差异大会形成稀疏 Doc Values。大量稀疏字段会浪费存储与遍历开销。建模上应尽量让同索引文档字段结构一致,不同业务的异构数据拆到独立索引。

6.3 聚合建模技巧

需求建模建议
按品牌统计brand 用 keyword
按时间统计date 字段,calendar_interval
多字段组合聚合预聚合宽表
大量唯一值聚合控制基数,必要时拆索引
{
  "aggs": {
    "by_brand": { "terms": { "field": "brand", "size": 20 } }
  }
}

7. Mapping 演进与重构

7.1 字段不可变与 reindex

字段类型创建后不可修改,调整分析器或类型必须重建索引。标准流程是新建索引、reindex、切换别名。

# 新建新版本索引
curl -X PUT 'http://localhost:9200/products_v2?pretty' \
  -H 'Content-Type: application/json' \
  -d '{"mappings": {"properties": {"name": {"type": "text", "analyzer": "ik_max_word"}}}}'

# 迁移数据
curl -X POST 'http://localhost:9200/_reindex?pretty' \
  -H 'Content-Type: application/json' \
  -d '{"source": {"index": "products_v1"}, "dest": {"index": "products_v2"}}'

7.2 别名切换实现零停机

curl -X POST 'http://localhost:9200/_aliases?pretty' \
  -H 'Content-Type: application/json' \
  -d '{
    "actions": [
      { "remove": { "index": "products_v1", "alias": "products" } },
      { "add":    { "index": "products_v2", "alias": "products" } }
    ]
  }'

应用层只访问别名 products,reindex 完成后原子切换,实现零停机迁移。alias 还支持按日期过滤与写索引分离。

7.3 reindex 的性能与注意

curl -X POST 'http://localhost:9200/_reindex?pretty' \
  -H 'Content-Type: application/json' \
  -d '{
    "source": { "index": "products_v1", "size": 5000 },
    "dest":   { "index": "products_v2" }
  }'

reindex 本质是滚动读取加批量写入,可在 source 用 query 过滤只迁移部分数据。大量数据迁移建议在低峰执行,并临时拉长 refresh 与关闭副本以提速。

7.4 Mapping 版本与文档规范

维护一份 mapping 变更记录,每个索引带版本号后缀(products_v1/v2)。删除旧索引前确认数据已备份且无查询引用,避免线上误删。

8. 总结

环节要点
Mapping 即表结构创建后字段难改,建模前置
类型选择text 全文、keyword 精确、date 时间、nested 保边界
dynamic生产用 false/strict,配合 dynamic_templates
多字段index/search analyzer 分离,multi-fields 按场景
嵌套选择频繁更新用 join,否则 nested 优先
聚合建模Doc Values 列存,字段结构尽量一致
演进重构reindex + 别名切换零停机
命名规范统一字段风格,控制字段数量

数据建模决定了 ES 的能力边界:检索靠倒排、排序靠 Doc Values、时间靠 date,mapping 一旦定型,后续优化成本极高。建模时应反复问自己每个字段的读法:要全文匹配、精确过滤、排序还是聚合。底部分词逻辑见《倒排索引与分词原理》,查询组合见《Query DSL 与相关性打分》。

延伸阅读

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「elasticsearch」更多文章

  1. 「搜索服务架构:从索引到容错」
  2. 「安全加固与访问控制:从角色到审计」
  3. 「地理空间搜索:从坐标到地图」