arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

接收者条件化潜在通信实现94%缓存回退

Receiver-Conditioned Latent Communication gives 94% CacheBack

Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang, Eugene Wu

arXiv 2609.32046首次发表:更新:

AI 中文总结

针对多智能体系统通信开销大的问题,提出接收者条件化通信方法CacheBack,通过接收者需求过滤压缩KV缓存,在FanOutQA上提升准确率14.7个百分点并降低3.2倍延迟。

AI 中文摘要

多智能体系统将大型上下文分布到多个智能体之间,这些智能体通过通信来协作解决任务。文本消息紧凑但需要解码,且可能遗漏接收智能体所需的证据。最近的潜在通信方法转而传输KV缓存,这避免了文本生成,并能提高准确性和降低延迟。然而,完整的KV缓存随单个智能体处理的上下文长度以及协同工作的智能体数量线性增长,这带来了内存和上下文成本,往往远超可用的GPU资源和上下文窗口大小。我们的关键观察是,智能体只需发送接收智能体执行其局部任务所需的信息——我们称之为接收者条件化通信。接收智能体向发送智能体传递其信息需求的简短描述,用于过滤和压缩发送智能体的KV缓存。CacheBack是一种简单、稳健、无需训练的接收者条件化实例,基于发送智能体的注意力权重实现。在FanOutQA基准上,使用Qwen 3的CacheBack移除了智能体原本会接收的75%的状态,将准确率提高了14.7个百分点,并将中位任务完成延迟相比文本通信降低了3.2倍。我们展示了在跨越密集Transformer、Mamba注意力混合模型和滑动窗口注意力的模型家族中具有可比的改进效果。

英文摘要

Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents need only send what the receiving agent requires for its local task -- which we call receiver-conditioned communication. The receiver agent passes the sender a small description of its information needs, which serves to filter and compress the sender agent's KV cache. CacheBack is a simple, robust, training-free instance of receiver conditioning based on the sender's attention weights. On FanOutQA, CacheBack with Qwen 3 removes 75% of the state the agent would otherwise receive, improving accuracy by 14.7 percentage points and reducing median task-completion latency by 3.2x relative to text communication. We show comparable improvements across model families that span dense Transformers, Mamba-attention hybrids, and sliding-window attention.

Comments29 pages, 13 figures, 13 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑