arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CacheRoute:面向大规模大语言模型服务的规划式前缀亲和路由

CacheRoute: Planned Prefix-Affinity Routing for Large-Scale LLM Serving

Huang Cheng

arXiv 2608.19677首次发表:更新:

发表机构

Meta(Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CacheRoute 是一种规划式前缀亲和路由方法,通过周期性路由规划解决大语言模型服务中前缀缓存复用与服务器过载的权衡,在 Llama-3.3-70B 模型上实现了更高 QPS 与 KV 缓存命中率。

AI 中文摘要

前缀缓存仅在重复请求返回仍保留前缀键值(KV)的服务器时才避免预填充。无缓存感知的负载均衡会分散这种复用,固定亲和性虽能保留复用却可能导致服务器过载。CacheRoute 通过周期性路由规划解决该权衡问题:它将高频率密钥纳入稳定的预热集,并根据预期负载分配其部署位置;热门密钥可使用多个目标,而我们的主要半合成聚合中的每个密钥仅使用一个。在 60 块 H100 GPU 上运行 fp8 精度的 Llama-3.3-70B 模型时,CacheRoute 在 3.5 秒 p99 服务水平目标(SLO)下可维持 176±11 的每秒查询数(QPS),是五个基准中最强者的 2.3 倍;服务的 KV 缓存命中率从无缓存感知负载均衡下的 64.1±1.3% 提升至 93.2±0.5%。第二个半合成聚合及受控的 8B 与突发实验分离了亲和性和部署位置的影响;两个 32B 工作负载提供了反例:当亲和性恢复的 KV 工作过少时,其残留负载偏差会降低或消除改进效果,因此我们建议通过影子回放对所有部署进行门控,而非仅根据工作负载统计启用亲和性。

英文摘要

Prefix caching avoids prefill only when a repeated request returns to a server that still holds the prefix KV. Cache-blind balancing disperses that reuse; fixed affinity preserves it but can overload a server. CacheRoute resolves this tradeoff with a periodic routing plan. It admits high-rate keys to a stable warm set and places their assignments by expected load. Hot keys may use more than one destination, although every key in our primary semi-synthetic aggregate uses exactly one. On Llama-3.3-70B in fp8 across 60 H100 GPUs, CacheRoute sustains 176+/-11 QPS at a 3.5-s p99 SLO, 2.3x the strongest of five baselines. Served KV-cache hit rate rises from 64.1+/-1.3% under cache-blind balancing to 93.2+/-0.5%. A second semi-synthetic aggregate and controlled 8B and burst experiments separate the effects of affinity and placement. Two 32B workloads provide the counterexamples: when affinity recovers too little KV work, its residual load skew reduces or erases the improvement. We therefore recommend gating any deployment with a shadow replay rather than enabling affinity from workload statistics alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑