arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SmartGen:支持无缝拆分式大语言模型推理的选择性KV缓存传输方案

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer

Xuchuan Luo, Jiacheng Shen, Xin Wang, Yangfan Zhou

arXiv 2607.28150首次发表:更新:

AI 中文总结

SmartGen是一款KV缓存传输引擎,通过三种数据传输路径实现拆分式LLM推理的KV缓存选择性传输,可将首token生成时间最多缩短4.3倍,且后续性能与准确率相当。

AI 中文摘要

当前大语言模型(LLM)服务系统普遍采用将LLM推理的预填充阶段和解码阶段拆分为两组独立节点的架构,但这种架构在自托管LLM部署于租赁云实例时面临重大挑战:拆分节点间传输海量键值(KV)缓存极易耗尽有限的节点间网络带宽。本文提出通过选择性传输关键KV缓存条目缓解网络瓶颈,实现选择性KV缓存传输需解决两大挑战:预填充阶段的精准KV选择,以及解码阶段的高效KV获取。为应对这些挑战,我们设计了SmartGen——一款支持三种数据传输路径的KV缓存传输引擎,可实现无缝拆分式LLM推理。具体而言,我们利用:1)基于配置文件的主动传输路径,在预填充阶段识别关键KV缓存条目并推送至解码节点;2)并行按需传输路径,在解码阶段同时获取远程和本地KV缓存条目;3)推测传输路径,最终将所有KV缓存交付至解码节点。实验结果表明,与典型的全KV缓存传输方案相比,SmartGen将首 token 生成时间最多缩短4.3倍,同时提供相当的后续解码性能和准确率。

英文摘要

Disaggregating the prefill and decoding stages of large language model (LLM) inference into two separate sets of nodes is widely adopted in today's LLM serving systems. However, such an architecture poses significant challenges for self-hosted LLM deployments on rented cloud instances, since transferring enormous key-value (KV) caches between disaggregated nodes can easily saturate the limited inter-node network bandwidth. In this paper, we propose to mitigate the network bottleneck by selectively transferring essential KV cache entries across the two stages. There are two challenges to achieve selective KV cache transfer, i.e., accurate KV selection during the prefill stage, and efficient KV fetching during the decoding stage. To address these challenges, we design SmartGen, a KV cache transfer engine that allows seamless disaggregated LLM inference with three data transfer paths. Specifically, we leverage 1) a profile-based proactive transfer path to identify and push essential KV cache entries to the decoding node during the prefill stage, 2) a parallel on-demand transfer path to simultaneously fetch remote and local KV cache entries during the decoding stage, and 3) a speculative transfer path to finally deliver all KV caches to the decoding node. Experimental results show that SmartGen reduces time-to-second-token by up to 4.3x compared with the typical full KV cache transfer approach while offering comparable subsequent decoding performance and accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑