<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Ai-Testing on PlumePHP</title><link>https://plumephp.com/tags/ai-testing/</link><description>Recent content in Ai-Testing on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Sat, 26 Sep 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/ai-testing/index.xml" rel="self" type="application/rss+xml"/><item><title>AI 模型测试实战：从 LLM 输出验证到 RAG 质量评估的全链路质量工程</title><link>https://plumephp.com/ai-model-testing/</link><pubDate>Sat, 26 Sep 2026 00:00:00 +0800</pubDate><guid>https://plumephp.com/ai-model-testing/</guid><description>&lt;p&gt;传统的测试方法论建立在&lt;strong&gt;确定性&lt;/strong&gt;之上：同样的输入必然得到同样的输出，断言 &lt;code&gt;assert x == y&lt;/code&gt; 是质量的黄金标准。但以 LLM 为代表的生成式 AI 模型打破了这一前提——同样的 prompt 每次调用可能返回不同的、且没有唯一正确答案的文本。&lt;code&gt;assert response == &amp;quot;Hello&amp;quot;&lt;/code&gt; 在 AI 测试中不再有意义，取而代之的是&amp;quot;这段回答是否语义正确？&amp;ldquo;&amp;ldquo;它是否忠实于给定的知识库？&amp;ldquo;&amp;ldquo;它是否包含了偏见？&amp;quot;。本指南系统覆盖 LLM 输出验证、RAG 质量评估、Prompt 回归测试、模型性能与对抗性测试，以及 MLOps CI/CD 中的模型质量门禁建设。&lt;/p&gt;</description></item><item><title>AI 辅助测试生成：LLM 时代的测试维护革命</title><link>https://plumephp.com/ai-assisted-testing/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0800</pubDate><guid>https://plumephp.com/ai-assisted-testing/</guid><description>&lt;p&gt;测试自动化领域长期面临一个尴尬的悖论：自动化测试本应减少人工投入，但编写和维护测试代码本身却消耗了与生产代码相当甚至更多的资源。研究表明，典型项目的测试代码量与生产代码量之比通常在 2:1 到 3:1 之间，而测试代码的维护成本往往高于生产代码——因为每次功能变更都可能导致大量测试用例的同步更新。大语言模型（LLM）的兴起为这一困境带来了颠覆性的解决方案。&lt;/p&gt;</description></item></channel></rss>