arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03893cs.LG

大语言模型家族中的跨模型KV缓存迁移:用于预填充复用的闭式线性映射

Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse

Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, Mohammad Mahdi Kamani, Ritika Borkar, Makesh Tarun Chandran, Pantea Zardoshti, Bita Darvish Rouhani

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型家族切换时需重新预填充的问题,提出跨模型KV缓存迁移方法,通过闭式线性映射器复用源模型KV缓存,可提升预填充速度,保留大部分准确率,具备实用性。

中文摘要 AI 辅助

生产部署中,常因成本-质量级联、对话中途切换及路由需求,在同一家族不同规模的模型间进行切换,每次切换都会迫使接收模型重新执行预填充。本文提出跨模型KV缓存迁移方法,让接收模型复用源模型的KV缓存,跳过预填充步骤。研究发现,当源模型与目标模型共享KV头数量及每头维度时,跨模型KV在匹配KV对上具有显著线性结构。在Qwen3 14B→32B场景中,单个源层可解释目标键56%的方差、值32%的方差,使用多个源层时该比例分别升至79%和65%。基于此,设计闭式岭回归映射器,按头操作,分三步实现:第一步,对每个目标层,选择预测性最高的前k个源层,将其KV拼接为输入;第二步,映射前去除键的旋转位置嵌入(RoPE),使拟合结果与位置无关,可跨上下文长度复用;第三步,在由500条长度为1024的FineWeb-Edu序列组成的小型校准集上拟合岭回归。令人惊讶的是,在三个家族的六组模型对中,该线性映射器在四组上保留了接收模型独立预填充准确率的73%-98%,而另外两组则大幅下降;非线性多层感知机(MLP)可在失效案例中恢复最多37个百分点的HellaSwag保留率。该映射器的运行速度比重做预填充快2.7至25倍,且在多轮交接时保持稳定,使跨模型KV缓存迁移具备实用性。

英文摘要

Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B->32B, one source layer explains 56% of variance in the target's keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-k most predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 73-98% of the receiver's standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +37 pp HellaSwag retention on the failures. The mapper runs 2.7-25x faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.

↑