用于跨模型KV共享的通用上下文复用层
A Universal Context-Reuse Layer for Cross-Model KV Sharing
浏览论文内容
中文总结 AI 辅助
该研究提出通用上下文复用层,实现跨模型KV共享,在同系列及跨系列LLM场景中降低预填充成本、提升准确率,验证了KV状态可作为可迁移计算表示。
中文摘要 AI 辅助
现代大语言模型(LLM)服务系统越来越多地在重复或共享的上下文上运行,但即便已有其他模型处理过相同输入,每个模型通常仍会执行自身的预填充计算。现有的KV缓存复用机制大幅减少了单一模型内的冗余计算,但通常假设缓存的生产者和使用者是相同的。我们研究跨模型KV共享,即把源模型产生的KV状态转换为可被不同目标模型使用的表示,这些目标模型可能在规模、架构、注意力配置、分词器及模型系列上存在差异。我们在同系列和跨系列设置中评估该方法:在Qwen2.5-7B→Qwen2.5-1.5B场景中,转换后的KV状态将LongBench2准确率从27.59%提升至34.48%,较原生1.5B基准提升6.89个百分点,同时降低了相对于原生目标预填充的切换成本;在跨系列Qwen2.5-1.5B→Gemma-2-2B场景中,4K上下文长度下KV切换可将目标侧预填充成本降低最多67.05%,同时解码困惑度接近原生模型基准;在更异构的Llama3.1-70B→Qwen2.5-7B场景中,跨系列切换实现了44.0%的准确率,而原生Qwen2.5-7B推理的准确率为45.7%,同时将实测延迟从899ms降至138ms。这些结果初步证明KV状态可作为可迁移的计算表示,而非严格的模型本地缓存,并推动上下文移动性作为一种系统抽象,用于减少异构LLM和多智能体推理工作流中的冗余预填充。
英文摘要
Modern large language model (LLM) serving systems increasingly operate over repeated or shared context, yet each model typically performs its own prefill computation even when another model has already processed the same input. Existing KV-cache reuse mechanisms substantially reduce redundant computation within a single model, but generally assume that the producer and consumer of a cache are identical. We study \emph{cross-model KV sharing}, which translates the KV state produced by a source model into a representation that can be consumed by a different target model, including models that differ in scale, architecture, attention configuration, tokenizer, and model family. We evaluate the approach in both within-family and cross-family settings. For Qwen2.5-7B $\rightarrow$ Qwen2.5-1.5B, translated KV states improve LongBench2 accuracy from 27.59\% to 34.48\%, a gain of 6.89 percentage points over the native 1.5B baseline, while reducing handoff cost relative to native target prefill. For the cross-family Qwen2.5-1.5B $\rightarrow$ Gemma-2-2B setting, KV handoff reduces target-side prefill cost by up to 67.05\% at 4K context length while maintaining decoding perplexity close to native-model baselines. In a more heterogeneous Llama3.1-70B $\rightarrow$ Qwen2.5-7B setting, cross-family handoff achieves 44.0\% accuracy compared with 45.7\% for native Qwen2.5-7B inference, while reducing measured latency from 899ms to 138ms. These results provide initial evidence that KV states can serve as transferable computational representations rather than strictly model-local caches, and motivate \emph{context mobility} as a systems abstraction for reducing redundant prefill across heterogeneous LLM and multi-agent inference workflows.
发表机构
- University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。