免预填充跨家族KV缓存迁移:面向异构多智能体大语言模型
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
浏览论文内容
中文总结 AI 辅助
针对异构多智能体LLM的跨家族KV缓存迁移问题,提出HeteroFold方法,通过结构对齐、缓存映射与校准实现免预填充传输,在多个基准上性能最优且速度大幅提升。
中文摘要 AI 辅助
近期多智能体大语言模型系统日益将异构模型组合用于专门的智能体角色。然而,基于文本的通信要求每个接收方对发送方已处理的共享上下文进行预填充。重用发送方的键值(KV)缓存可避免这种冗余,但跨模型家族的免预填充迁移必须处理分词、模型深度和KV表示上的差异。为解决这些问题,我们提出HeteroFold,一种免预填充的跨家族KV缓存迁移方法,该方法保持发送方和接收方均冻结。HeteroFold对齐模型结构,将发送方缓存映射到接收方空间,并对其进行校准以保持接收方行为。在六个迁移方向上,HeteroFold在所有四个长上下文基准和大多数短上下文设置中取得了最佳的缓存迁移性能。它还在多智能体基准上与基于文本的通信相匹配。在32K上下文长度下,Llama-3.1-8B到Ministral-3-14B的迁移比原生预填充快10.7倍,比最先进的免预填充基线Dense Latent和KV Ridge快1.18至1.47倍。这些结果表明,HeteroFold无需接收方预填充即可实现高效的跨家族KV重用。
英文摘要
Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structures, maps the sender cache into the receiver space, and calibrates it to preserve receiver behavior. Across six transfer directions, HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings. It also matches text-based communication on the multi-agent benchmark. At 32K context length, Llama-3.1-8B$\rightarrow$Ministral-3-14B transfer is $10.7\times$ faster than Native Prefill and $1.18$--$1.47\times$ faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge. These results show that HeteroFold enables efficient cross-family KV reuse without receiver prefill.