大语言模型家族中的跨模型KV缓存迁移:用于预填充复用的闭式线性映射
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear Mapping for Prefill Reuse
浏览论文内容
中文总结 AI 辅助
针对大语言模型家族切换时需重新预填充的问题,提出跨模型KV缓存迁移方法,通过闭式线性映射器复用源模型KV缓存,可提升预填充速度,保留大部分准确率,具备实用性。
中文摘要 AI 辅助
生产部署中,常因成本-质量级联、对话中途切换及路由需求,在同一家族不同规模的模型间进行切换,每次切换都会迫使接收模型重新执行预填充。本文提出跨模型KV缓存迁移方法,让接收模型复用源模型的KV缓存,跳过预填充步骤。研究发现,当源模型与目标模型共享KV头数量及每头维度时,跨模型KV在匹配KV对上具有显著线性结构。在Qwen3 14B→32B场景中,单个源层可解释目标键56%的方差、值32%的方差,使用多个源层时该比例分别升至79%和65%。基于此,设计闭式岭回归映射器,按头操作,分三步实现:第一步,对每个目标层,选择预测性最高的前k个源层,将其KV拼接为输入;第二步,映射前去除键的旋转位置嵌入(RoPE),使拟合结果与位置无关,可跨上下文长度复用;第三步,在由500条长度为1024的FineWeb-Edu序列组成的小型校准集上拟合岭回归。令人惊讶的是,在三个家族的六组模型对中,该线性映射器在四组上保留了接收模型独立预填充准确率的73%-98%,而另外两组则大幅下降;非线性多层感知机(MLP)可在失效案例中恢复最多37个百分点的HellaSwag保留率。该映射器的运行速度比重做预填充快2.7至25倍,且在多轮交接时保持稳定,使跨模型KV缓存迁移具备实用性。
英文摘要
Production deployments often swap between different-sized models in a family for cost-quality cascading, mid-conversation switching, and routing, and each swap forces the receiver to repay the prefill from scratch. We propose cross-model KV cache transfer, where the receiver reuses the source's KV cache, skipping prefill. We find that cross-model KV has substantial linear structure across matched-KV pairs, where source and target share KV head count and per-head dimension. On Qwen3 14B->32B, one source layer explains 56% of variance in the target's keys and 32% in values, rising to 79% and 65% with multiple source layers. Building on this, we design a closed-form ridge mapper that operates per head and proceeds in three steps. First, for each target layer we select the top-k most predictive source layers and concatenate their KV as input. Second, we strip RoPE from the keys before mapping, so the fit is position-free and reusable across context lengths. Third, we fit ridge regression on a small calibration set of 500 FineWeb-Edu sequences of 1,024 tokens each. Surprisingly, across six pairs in three families, this linear mapper retains 73-98% of the receiver's standalone-prefill accuracy on four pairs, while two degrade sharply. A nonlinear MLP recovers up to +37 pp HellaSwag retention on the failures. The mapper runs 2.7-25x faster than re-prefill and remains stable across multi-turn handoff, making cross-model KV cache transfer practical.