arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20617cs.AIcs.LG

异构语言模型间的双缓存隐空间通信

Dual-Cache Latent Space Communication between Heterogeneous Language Models

Jiyao Liu, Qi Zhang, Yaoyi Jia, Ziwen Kan, Song Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出XKV协议,解决异构语言模型隐空间通信的三个限制,仅训练转换器,在45个数据集-模型对设置中,其性能、速度均优于现有方法,参数效率更高。

中文摘要 AI 辅助

多智能体大语言模型(LLM)系统在不同模型间分配工作,因此回答问题往往需要其他智能体上下文中的知识:共享器(Sharer)编码了接收器(Receiver)完成任务所需的信息。这类系统通常通过交换文本通信,这会让自回归解码处于关键路径,且通信内容是在未考虑接收器状态的情况下生成的离散消息。近期的隐空间协议则将共享器的键值(KV)缓存转换为接收器的缓存:C2C支持异构模型,但要求两者读取相同输入;LCF-X则通过无位置的共享器缓存池化消除了共享上下文的要求。不过仍存在三个限制:LCF-X仅压缩共享器,向所有接收器位置提供相同的层局部摘要,且假设两者层数和KV几何匹配。本文提出XKV解决这三个问题:用学习查询注意力池化两个缓存;通过与接收器对齐的层标记的自注意力,结合学习到的层映射协调不同深度,将池化摘要混合为紧凑的联合内存;共享位置解码器让每个原始接收器缓存位置能在接收器原生KV几何中检索自身的每头门控残差。所有模型保持冻结,可在系列、深度、KV头数、头维度、分词器上存在差异,仅训练转换器。在45个数据集-模型对设置(6个异构配对、3个同模型有序配对、5个数据集)中,XKV取得最高宏观分数和最佳平均排名,在所有数据集上均优于LCF-X(在ROPES数据集上精确匹配提升4.6,F1提升4.2),在5个数据集中的4个优于文本通信;训练参数减少76%,缓存对翻译速度提升10.3倍(5.8毫秒对比59.9毫秒);端到端速度比LCF-X快26%,比文本通信快6.8倍。

英文摘要

Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has encoded information that a Receiver needs to complete its task. They usually communicate by exchanging text, which puts autoregressive decoding on the critical path and reduces the exchange to a discrete message written without sight of the receiver's state. Recent latent protocols instead translate the sharer's key-value (KV) cache into the receiver's: C2C supports heterogeneous models but requires both to read the same input, while LCF-X removes this shared-context requirement through position-free sharer-cache pooling. Three restrictions remain: LCF-X compresses the sharer alone, supplies the same layer-local summary to every receiver position with no joint cross-layer memory to retrieve from, and assumes matched layer count and KV geometry. We introduce XKV, which lifts all three: learned-query attention pools both caches; self-attention over receiver-aligned layer tokens, with a learned layer map reconciling different depths, mixes the pooled summaries into a compact joint memory; and a shared position decoder lets every raw receiver cache position retrieve its own per-head-gated residual in the receiver's native KV geometry. Both models stay frozen and may differ in family, depth, KV-head count, head dimension, and tokenizer; only the translator is trained. Across 45 dataset-model-pair settings (six heterogeneous and three same-model ordered pairings, five datasets), XKV attains the highest macro score and best average rank, improving on LCF-X on every dataset (by 4.6 exact-match and 4.2 F1 points on ROPES) and surpassing text communication on four of the five, while training 76% fewer parameters and translating a cache pair 10.3x faster (5.8 vs. 59.9 ms); end to end, XKV is 26% faster than LCF-X and 6.8x faster than text communication.

发表机构

  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

↑