<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Levenshtein on PlumePHP</title><link>https://plumephp.com/tags/levenshtein/</link><description>Recent content in Levenshtein on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Mon, 28 Sep 2026 10:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/levenshtein/index.xml" rel="self" type="application/rss+xml"/><item><title>文本相似度与模糊匹配：Levenshtein、Jaro-Winkler 与 SimHash</title><link>https://plumephp.com/others-fuzzy-text-matching/</link><pubDate>Mon, 28 Sep 2026 10:00:00 +0800</pubDate><guid>https://plumephp.com/others-fuzzy-text-matching/</guid><description>&lt;h2 id="引言"&gt;引言&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;用户把 &lt;code&gt;brother&lt;/code&gt; 打成 &lt;code&gt;borther&lt;/code&gt;，客服把客户名抄错一位，爬虫抓到了 99% 相同的两篇文章&amp;rdquo;——&lt;strong&gt;模糊匹配&lt;/strong&gt;要回答的就是&amp;quot;这两段文本像不像、差几个字符、是不是同一个东西&amp;quot;。本文从零搭一套匹配工具箱：先讲最经典的&lt;strong&gt;编辑距离&lt;/strong&gt;（Levenshtein 及其优化、Damerau 的调换），再讲 token 级的 &lt;strong&gt;Jaccard/Dice&lt;/strong&gt; 与擅长人名匹配的 &lt;strong&gt;Jaro-Winkler&lt;/strong&gt;，接着讲模糊搜索怎么落地（fzf 原理、编辑距离阈值、候选集），再讲&lt;strong&gt;大规模去重&lt;/strong&gt;的利器 SimHash/MinHash（为什么不是两两比较），最后给一个可落地的&amp;quot;匹配引擎&amp;quot;设计。&lt;/p&gt;</description></item></channel></rss>