arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00891cs.AI

CacheBridge:高效的跨模型KV缓存迁移

CacheBridge: Efficient Cross-Model KV Cache Transfer

Xingyu Qu, Siyuan Lu, Zhiyu Chen, Sheng Wang, Tao Lin

首次发表
浏览论文内容

中文总结 AI 辅助

CacheBridge是一种高效的跨模型KV缓存迁移方法,通过架构索引映射器等设计,在Qwen3等模型上实现了存储、速度与精度的优化,解决了全头映射的缺陷。

中文摘要 AI 辅助

在多模型系统中,大型语言模型(LLM)之间共享上下文时,由于KV缓存具有模型特异性,接收模型需要对共享前缀进行预填充。近期的闭式跨模型KV迁移方法(即全头映射,Full-Head Mapping)通过拟合一种无需训练的仿射映射器来实现源缓存到目标缓存的转换,从而避免了预填充过程。然而,其全头设计会从所选层的每个源KV头映射每个目标KV头,导致迁移质量对架构差异敏感,且映射器的存储和应用成本随支持层数增加而增长。为此,本文提出CacheBridge,该方法协同设计了架构索引的映射器支持、注意力对齐的校准以及有界映射器构建,同时保留了用于在线部署的闭式仿射接口。CacheBridge限制每个目标头仅映射至匹配的源头,通过因果注意力敏感性对重构误差加权,并使用融合GPU内核构建加权充分统计量,无需显式存储完整观测张量。在三种迁移方向上,CacheBridge在全头映射(Full-Head Mapping)损失大量精度的两个Ministral 3迁移方向上恢复了精度,同时在Qwen3上保持了99.83%的平均目标保留率。在Qwen3 14B→32B的迁移中,它将映射器存储减少了8倍,应用速度最高提升3.0倍,仅用十分之一的校准数据即可达到全头映射的性能,且将500个序列的构建时间从92.63秒缩短至8.63秒(加速10.7倍)。

英文摘要

Sharing context between LLMs in a multi-model system requires the receiving model to prefill the shared prefix because KV caches are model-specific. Recent closed-form cross-model KV transfer, hereafter Full-Head Mapping, avoids this replay by fitting a training-free affine mapper from source to target caches. However, its full-head design maps each target KV head from every source KV head in the selected layers, making transfer quality sensitive to architectural differences and causing mapper storage and application cost to grow with layer support. To this end, we introduce CacheBridge, which co-designs architecture-indexed mapper support, attention-aligned calibration, and bounded mapper construction while retaining a closed-form affine interface for online deployment. CacheBridge restricts each target head to a matched source head, weights reconstruction errors by causal attention sensitivity, and uses a fused GPU kernel to construct weighted sufficient statistics without materializing full observation tensors. Across three transfer directions, CacheBridge recovers the two Ministral 3 transfer directions where Full-Head Mapping loses substantial accuracy while preserving 99.83\% mean target retention on Qwen3. On Qwen3 $14\mathrm{B}\to32\mathrm{B}$, it reduces mapper storage by $8\times$, accelerates application by up to $3.0\times$, matches \fullhead with one tenth of the calibration data, and reduces 500-sequence construction from 92.63 to 8.63 seconds ($10.7\times$).

发表机构

  • Westlake University(西湖大学)
  • Wuhan University(武汉大学)
  • Amazon(亚马逊公司)

机构由 AI 辅助整理,请以论文原文为准。

↑