Clojure 数据科学与数值计算:tablecloth / tech.ml 实战

深入 Clojure 生态的数据科学实践:tablecloth 数据框操作与 pipeline 处理、tech.ml 的 dataset/column 抽象、数值计算与统计分析、数据可视化、机器学习建模与模型评估,展示函数式风格如何让数据清洗与分析代码更简洁可复现。

数据分析常常被等同于 Python 的 pandas + sklearn,但 Clojure 在数据处理上其实有着独特的优势:不可变数据天然可复现、REPL 驱动让探索式分析极其顺畅、->> 线程宏让 pipeline 式清洗一目了然。本文用 tablecloth(Clojure 的 pandas 等价物)与 tech.ml(机器学习平台)走通一条从数据加载、清洗、可视化到建模评估的完整数据科学链路。

1. Clojure 数据科学生态地图

1.1 核心库选型

库定位对标 Python
tablecloth数据框(DataFrame)操作pandas
tech.ml.dataset底层 dataset/column 抽象numpy + pandas 内核
tech.ml机器学习工作流sklearn
incanter统计分析 + 绘图(较老)scipy / matplotlib
hanami / nvgVega-Lite 声明式可视化altair
libpython-clj调用 Python 生态互操作

组合建议:数据处理用 tech.ml.dataset + tablecloth,建模用 tech.ml,可视化用 hanami(Vega-Lite)。它们共享 tech.ml.dataset 的数据结构,可以无缝衔接。

2. 用 tablecloth 做数据清洗

2.1 加载数据

(require '[tablecloth.api :as tc])

;; 从 CSV 加载成数据框(ds = dataset)
(def df
  (tc/dataset "orders.csv"
              {:parser-fn {:order_id :integer
                           :amount    :double
                           :paid_at   :datetime}}))

(tc/head df)  ; 查看前几行
;; => #tech.ml.dataset/table [...](类似 pandas 的 df.head())

2.2 pipeline 式清洗

tablecloth 的 API 围绕 ->> 设计,清洗过程是一条清晰的流水线:

(def clean-df
  (-> df
      ;; 按列处理:缺失值填充、类型转换
      (tc/replace-missing :amount {:value 0})
      (tc/convert-types :amount :double)
      ;; 筛选:只保留近一年的有效订单
      (tc/select-rows (fn [row] (and (:paid_at row)
                                     (> (:amount row) 0))))
      ;; 新增列:月与金额档位
      (tc/add-column :month #(month (:paid_at %)))
      (tc/add-column :amount-tier
                     #(cond (< (:amount %) 100) :low
                            (< (:amount %) 500) :mid
                            :else :high))
      ;; 排序 + 去重
      (tc/order-by [:paid_at :desc])
      (tc/distinct-by :order_id)))

关键:每一步都返回新的 dataset(不可变),中间任何一步出问题都可以回到 REPL 检查该步的中间结果——这就是「可复现数据分析」。

2.3 分组聚合

;; 按月份 + 金额档位做汇总:group-by + aggregate
(def monthly
  (-> clean-df
      (tc/group-by [:month :amount-tier])
      (tc/aggregate {:order-count #(count %)
                     :total-amount #(reduce + (map :amount %))})))

;; 与 pandas 的 groupby().agg() 完全对标

3. 数值计算与统计分析

3.1 统计量计算

(require '[tech.v3.datatype.functions :as fns])
(require '[tech.ml.dataset.column :as col])

;; 对一列做描述性统计
(def amounts (col/->array (:amount clean-df)))

{:mean (fns/mean amounts)
 :std  (fns/std amounts)
 :min  (fns/min amounts)
 :p90  (fns/quantile amounts 0.9)}

3.2 相关性分析

;; 用 tech.ml.dataset 计算两列相关性
(defn corr [ds a b]
  (let [xs (col/->array (a ds))
        ys (col/->array (b ds))]
    (fns/correlation xs ys)))

(corr clean-df :amount :order-count)

4. 数据可视化:hanami + Vega-Lite

4.1 声明式绘图

(require '[hanami.core :as h])
(require '[tablecloth.api :as tc])

;; 柱状图:每月订单量
(def bar-spec
  (h/xform :bar-chart
           {:data  (-> monthly (tc/rows))
            :x     :month
            :y     :order-count}))

;; hanami 生成 Vega-Lite 规范(JSON),可在浏览器渲染
(println bar-spec)

4.2 为什么声明式

Vega-Lite 的声明式规范(data + mark + encoding)与 Clojure 的「数据即代码」精神高度契合——图表本身就是数据,可组合、可版本化、可复用。改一个 encoding 字段就是一张新图。

5. 机器学习建模:tech.ml

5.1 数据切分

(require '[tech.ml :as ml])

;; 把 dataset 拆成特征与标签
(def split-data
  (ml/feature-dataset->row-major-feature-dataset
    clean-df
    {:feature-cols [:amount :order-count :month]
     :target-cols  [:amount-tier]}))

;; 训练 / 测试集划分
(def [train test] (ml/split-dataset split-data {:seed 42 :split-percent 0.8}))

5.2 训练与评估

;; 训练一个随机森林分类器
(def trained
  (ml/train {:model-type :classification/random-forest
             :dataset    train}))

;; 在测试集上评估
(def report
  (ml/evaluate {:trained-model trained
                :dataset       test
                :metric-map    {:accuracy :classification/accuracy
                                :f1       :classification/f1}}))

(report :accuracy)
;; => 0.87

5.3 模型持久化

;; 保存 / 加载模型,用于服务端推理
(ml/save-model trained "models/order-tier.model")
(def reloaded (ml/load-model "models/order-tier.model"))

6. 与 Python 生态互操作

6.1 libpython-clj

当 Clojure 生态缺某个算法时,可以用 libpython-clj 直接调用 Python 库:

(require '[libpython-clj2.require :refer [require-python]])
(require '[libpython-clj2.python :as py])

(require-python '[numpy :as np])
(require-python '[sklearn.linear_model :as lm])

(def X (py/py-tuple->list (np/array (map #(vector (:amount %) (:order-count %)) (tc/rows clean-df)))))
(def model (lm/LinearRegression.))
(py/call-attr model :fit X (py/py-tuple->list (map :month (tc/rows clean-df))))

判断:能纯 Clojure 解决的(数据处理、常规建模、可视化)优先 Clojure;遇到深度神经网络等生态壁垒时再互操作 Python——两条路径在同一 REPL 里协作。

7. 数据科学工程化

7.1 可复现性

实践说明
固定 seed划分/采样随机种子固定,结果可复现
数据版本化用 deps.edn + git 管理数据与代码
pipeline 缓存清洗结果缓存,避免重复计算
全量测试用属性测试(/clojure-property-testing/)验证清洗逻辑

7.2 与分析/生产的边界

探索分析(REPL)与生产服务是两个世界:

  • 探索:REPL 里随意玩,聚焦「发现问题」
  • 生产:把清洗 pipeline 固化为函数,用 /clojure-modern-web-stack/ 的 Ring 服务包装成模型 API
  • 监控:模型上线后要跟踪输入分布漂移与预测质量(数据漂移检测是数据科学工程的关键一环)

8. 常见陷阱

陷阱现象规避
大 DataFrame 性能全量拷贝开销大用 tech.ml 底层数组、减少整表变换
内存爆炸惰性 seq 与 dataset 混用明确 realize,控制物化点
日期/时区混乱聚合错月统一 UTC 存储、展示时转换
混淆数据框与 seqAPI 调用错位先 (tc/rows df) 再操作 seq

9. 总结

Clojure 数据科学栈(tech.ml.dataset + tablecloth + tech.ml + hanami)覆盖了从数据加载、清洗、聚合、可视化到建模评估的完整链路,且因为不可变与 REPL 驱动,天然具备可复现性——这是 pandas 流水线很难企及的优势。数据清洗用 ->> 写 pipeline、建模用 tech.ml 的 train/evaluate、可视化用 Vega-Lite 声明式规范,需要深度模型时再互操作 Python。结合 /clojure-data-pipeline/ 的流处理与 /clojure-testing-quality/ 的质量保障,Clojure 完全可以作为数据分析与模型服务的主战场。

延伸阅读

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「clojure」更多文章

  1. Clojure 函数式错误处理:Result、异常与结构化错误
  2. Clojure GraphQL API 实战:lacinia、Schema、Resolver 与权限
  3. Clojure REPL 驱动开发:nREPL、热重载与交互式工作流