<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>TensorRT-LLM on PlumePHP</title><link>https://plumephp.com/tags/tensorrt-llm/</link><description>Recent content in TensorRT-LLM on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Thu, 10 Sep 2026 10:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/tensorrt-llm/index.xml" rel="self" type="application/rss+xml"/><item><title>TensorRT-LLM：大语言模型的高效推理部署</title><link>https://plumephp.com/ai-tensorrt-llm/</link><pubDate>Thu, 10 Sep 2026 10:00:00 +0800</pubDate><guid>https://plumephp.com/ai-tensorrt-llm/</guid><description>&lt;p&gt;大语言模型（LLM）的推理部署是 AI 基础设施的核心挑战之一。随着模型参数量从数十亿增长到数千亿，如何在保证低延迟和高吞吐的前提下高效运行这些模型，成为生产环境的必答题。NVIDIA 推出的 &lt;strong&gt;TensorRT-LLM&lt;/strong&gt; 正是为了解决这一痛点而诞生的专用推理框架。本文将系统性地介绍 TensorRT-LLM 的核心机制、优化手段和落地实践。&lt;/p&gt;</description></item></channel></rss>