arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

混合翻译器:跨异构大语言模型的KV缓存转换

KV Cache Translation across Heterogeneous Large Language Models

Jin-woo Lee, Minkyung Song, Junghyun Oh, Seunghoon Han, Gwangseon Jang, Soyoung Park, Sungsu Lim

arXiv 2607.28979首次发表:更新:

AI 中文总结

提出MoT框架解决异构LLM间KV缓存无法复用问题,通过多翻译器模块和上下文校正损失实现缓存转换,在多模型转换和实际场景中验证了性能与复用效果。

AI 中文摘要

异构大语言模型(LLM)系统日益依赖共享上下文、检索证据和多智能体对话历史,但其内部键值(KV)缓存仍为模型专属,无法跨架构复用。因此,每个模型必须重复预填充或存储相同上下文的缓存,限制了多模型推理和长上下文生成的可扩展性。我们提出混合翻译器(Mixture-of-Translators,MoT),一种将源LLM的上下文KV缓存映射到目标LLM缓存空间的缓存转换框架。与依赖单一投影路径或全局共享潜在空间的现有方法不同,MoT使用多个翻译器模块捕获多样化的源-目标映射。为进一步降低残余转换误差,我们引入上下文校正损失,使重放的目标轨迹与原生目标轨迹对齐。我们揭示了缓存转换中的两种相互竞争的失效模式:早期注入导致的传播转换偏移和晚期注入导致的末态偏移。MoT通过翻译器混合和目标端校正解决这些问题。在Qwen2.5、GPT-2和OPT模型的同构与异构转换中,MoT保留了下游问答(QA)性能,包括Qwen2.5-7B规模转换的51.0%平均闭集QA准确率和0.43平均抽取式QA F1。在实际案例研究中,MoT实现了多智能体推理的保质量内存复用,并在长上下文缓存增强生成中保留了96.3%的直接上下文质量,证明了异构LLM间可扩展的KV缓存复用能力。

英文摘要

Heterogeneous Large Language Model (LLM) systems increasingly share contexts, retrieved evidence, and multi-agent dialogue histories, yet their internal key-value (KV) caches remain model-specific and cannot be reused across architectures. Consequently, each model must repeatedly prefill or store caches for the same context, limiting the scalability of multi-model reasoning and long-context generation. We propose Mixture-of-Translators (MoT), a cache translation framework that maps context KV caches from a source LLM into the cache space of a target LLM. Unlike prior approaches that depend on a single projection path or global shared latent space, MoT uses multiple translator modules with token-level routing to capture diverse cache translations. We further introduce a Context Correction Loss that aligns the replayed target trajectory with the native target trajectory, reducing residual translation error. Our analysis reveals two competing failure modes: propagation error from early injection and correction-deficit error from late injection. We evaluate heterogeneous translation among Llama, Gemma, and Qwen, spanning substantially different architectures and KV-cache spaces. MoT achieves an average accuracy of 57.6% and F1 of 0.42. These scores recover 95.0% and 120.6% of native target performance and reach 1.5x and 3.5x the strongest-baseline scores, respectively. In practical case studies, MoT enables accurate heterogeneous-agent discussion with scale-invariant KV-memory behavior as the number of agents grows, while retaining near-native generation quality in long-context cache-augmented generation. These results demonstrate scalable KV-cache reuse across heterogeneous LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑