arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KV-Lingo:通过蒸馏学习 KV 缓存翻译器

KV-Lingo: Learning KV-Cache Translators with Distillation

Valérie Castin, Keitaro Sakamoto, Anastasiia Filippova, João Monteiro, Marco Cuturi, Pierre Ablin

arXiv 2609.32610首次发表:更新:

AI 中文总结

KV-Lingo 通过蒸馏学习线性映射,将源模型的 KV 缓存翻译为目标模型可读形式,避免模型切换时重新预填充,显著降低首 token 延迟,支持高效动态模型路由与无缝切换。

AI 中文摘要

大型语言模型通过键值(KV)缓存来表示上下文。缓存是模型特定的:对于相同的文本,具有不同架构或权重的模型会产生不兼容的表示。这使得在共享上下文上切换模型代价高昂:尽管上下文已经被一个模型处理过,但新来的模型必须再次处理它才能构建自己的缓存。我们引入了 KV-Lingo,一种将源模型的 KV 缓存翻译为目标模型可读取缓存的方法。KV-Lingo 由一组线性映射组成,通常目标模型的每一层对应一个映射,这些映射独立应用于所有 token 的键和值表示。我们使用蒸馏来训练这些映射,最小化目标模型从其原生缓存和翻译缓存进行预测之间的差异。我们考虑了多对跨越多种规模和架构的模型,在通用文本语料库上为每对模型训练一个翻译器。所得到的翻译器在小到大规模和大到小规模的迁移中均保持了强大的下游性能。由于切换现在只需一个线性映射和一次解码步骤,而不是预填充,用缓存翻译替代重新预填充,在 Apple M3 Ultra 上对 Qwen 模型进行 64 token 提示的模型切换后,首 token 时间减少了 9.6 倍,在 H100 上 32k 上下文长度时最多减少 29 倍。这些优势使 KV-Lingo 在动态模型路由中特别有用:上下文可以由一个模型处理,仅在需要时交给另一个模型,而无需重新预填充共享前缀。最后,我们展示了 KV-Lingo 可用于无缝模型切换,在我们的多轮评估中,在重复切换时保持接近重新预填充的性能。

英文摘要

Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been processed by one model, the incoming model must process it again to build its own cache. We introduce KV-Lingo, a method for translating the KV cache of a source model into one that can be read by a target model. KV-Lingo consists of a collection of linear maps, typically one per layer of the target model, that are applied independently on all tokens' key and value representations. We train these maps using distillation, minimising the divergence between the target model's predictions from its native cache and those from the translated cache. We consider several model pairs spanning multiple sizes and architectures, training one translator per pair on a generic text corpus. The resulting translators preserve strong downstream performance in both small-to-large and large-to-small transfers. Since a switch then costs a linear map and a single decoding step instead of a prefill, replacing re-prefill with cache translation reduces the time to first token after a model switch by 9.6x already on a 64-token prompt for Qwen models on an Apple M3 Ultra, and by up to 29x at 32k context length on an H100. These gains make KV-Lingo particularly useful for dynamic model routing: a context can be processed by one model and handed off to another only when needed, without re-prefilling the shared prefix. We finally show that KV-Lingo can be used for seamless model switching, staying close to re-prefill across repeated switches in our multi-turn evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑