AI 中文总结
针对异构计算与传输成本下LLM前缀KV放置的次优问题,提出PrefixPlace规划器,其性能接近最优且效率高,可提升RAG等场景的成本节省效果。
AI 中文摘要
前缀键值(KV)复用可避免大语言模型(LLM)推理中的重复预填充,但本地缺失会导致重新计算或副本获取。其相对成本随硬件、前缀深度、KV吞吐量及副本位置变化,基于命中率的放置方案并非最优。为解决该问题,我们提出epoch级规划器PrefixPlace,其在内存预算及已分析的需求、计算、传输成本下分配前缀完整目标。目标分解为本地副本价值加首副本覆盖率,且依赖源的成本产生单调设施定位目标;每个工作节点更新为可在O(nk)时间精确求解的加性根树问题(n为块数,k为容量),得到固定阶1/2近似解,且通过坐标优化和序多样起始可进一步改进。基于T4、L4和A100的测量结果显示出不同运行模式:在432个含精确最优解的实例中,PrefixPlace平均达到最优解的99.84%,且从未低于98.02%;在检索增强生成(RAG)回放中,其物化成本节省较vLLM自动前缀缓存(vLLM-APC)提升40.3%,较最佳离线基准提升6.3%;在WikiQA上,该提升分别为40.4%和5.3%。最后,PrefixPlace可在单个处理器上用12.3秒解决50000个节点、16个工作节点的放置问题,支持及时重规划。
英文摘要
Prefix Key-Value (KV) reuse avoids repeated prefill in Large Language Model (LLM) inference, but local misses require recomputation or replica fetches. Their relative cost varies with hardware, prefix depth, KV goodput, and replica location, making hit-rate-based placement suboptimal. To address this issue, we propose an epoch-level planner, PrefixPlace, which assigns prefix-complete targets under memory budgets and profiled demand, compute, and transfer costs. The objective decomposes into local-copy value plus first-replica coverage, and source-dependent costs yield a monotone facility-location objective; each worker update is an additive rooted-tree problem solved exactly in O(nk) time for n chunks and capacity k, giving a fixed-order 1/2-approximation that coordinate refinement and order-diverse starts improve without weakening. T4, L4, and A100 measurements reveal distinct regimes. Across 432 instances with exact optima, PrefixPlace averages 99.84% of optimum and never falls below 98.02%. In Retrieval-Augmented Generation (RAG) replays, it improves materialization-cost saving by 40.3% over vLLM Automatic Prefix Caching (vLLM-APC) and 6.3% over the best offline baseline. On WikiQA, gains are 40.4% and 5.3%. Finally, PrefixPlace solves a 50,000-node, 16-worker placement in 12.3 s on one processor, enabling timely replanning.