Repository navigation
p/llm-terminology-10-inference-serving/ #1
Replies: 1 comment
|
Hi. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
p/llm-terminology-10-inference-serving/
这是「大模型训练与推理术语全景图」系列的第 10 篇,覆盖推理服务与性能指标的 6 个概念:Prefill/Decode、TTFT、TPOT、Prefix Caching / RadixAttention、Chunked Prefill、Disaggregated Inference。 主要参考框架/服务:vLLM、SGLang、TGI、DeepSeek 推理服务、Sarathi-Serve、DistServe、Splitwise\n开篇:推理服务的两个阶段和两个指标 LLM 推理不是一次"模型前向",而是两个特性截然不同的阶段。Prefill(预填充)一次性处理整个 prompt,构造 N×N 注意力矩阵、打满 GPU 算力,是 compute-bound;Decode(解码)逐 token 生成,每步都要读回全部模型权重与 KV Cache、算力利用率往往低于 5%,是 memory-bound。用户感知到的两个维度恰好对应这两阶段:TTFT(首 token 延迟)由 prefill 主导,决定"是不是即时响应";TPOT(单 token 生成时间)由 decode 主导,决定"是不是流畅"。高效推理服务(vLLM、SGLang)的所有优化都围绕这两阶段展开——Chunked Prefill 把长 prompt 切块,避免 prefill 阻塞正在 decode 的请求;Prefix Caching 复用历史 KV,多轮对话 TTFT 降低 50-90%;Disaggregated Inference 干脆把两阶段分池部署,prefill 池高算力、decode 池高带宽,吞吐再提 30-60%。本篇按"两阶段 → 两指标 → 三类优化"的脉络梳理这 6 个术语。\n推理服务与性能指标(Inference Serving & Metrics) LLM 推理不是简单的"模型前向"。高效推理服务需要理解 prefill/decode 两阶段、关键延迟指标和缓存策略。\nPrefill/Decode 缩写:—\n
https://bevisy.github.io/p/llm-terminology-10-inference-serving/
All reactions