arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向任务的关键层KV通信用于高效潜在多智能体协作

Task-Oriented Key-Layer KV Communication for Efficient Latent Multi-Agent Collaboration

Dongsen Zhang, Peipei Li, Zekun Li, Wenjun Xu

arXiv 2610.08820首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; University of California, Santa Barbara(北京邮电大学; 加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在多智能体通信开销大、冗余多的问题,提出无需训练的KITE框架,通过接收方轨迹失真准则选择关键层并仅传输该层工作记忆,实现通信量减少28-36倍、推理加速3倍且准确率提升23.3个百分点。

AI 中文摘要

基于大型语言模型的多智能体系统通过协作提升复杂问题求解能力,而潜在通信直接传输模型内部状态以避免自然语言的高推理成本。然而,现有的基于KV的潜在通信方法优先考虑发送方状态保真度,导致大量通信和计算开销,并可能引入冗余信息。为解决这些局限,我们从任务导向的角度重新审视潜在通信,将其目标从发送方状态保真度转变为接收方任务充分性。在此框架下,我们提出KITE,一种无需训练的面向任务的关键层KV通信框架。KITE利用接收方轨迹失真准则识别任务有效的关键层,仅传输与该关键层关联的潜在工作记忆,并进一步使用同一层作为自回归潜在推理的入口点。在七个基准、两个模型家族和三种模型规模上的实验表明,与全层KV通信相比,KITE将通信量减少28-36倍,实现高达3倍的端到端推理加速,并将准确率提升高达23.3个百分点。

英文摘要

Large language model-based multi-agent systems improve complex problem solving through collaboration, while latent communication directly transmits model internal states to avoid the high inference costs of natural language. However, existing KV-based latent communication methods prioritize sender-side state fidelity, leading to substantial communication and computation overhead and potentially introducing redundant information. To address these limitations, we revisit latent communication from a task-oriented perspective, shifting its objective from sender-side state fidelity to receiver-side task sufficiency. Under this formulation, we propose KITE, a training-free framework for task-oriented key-layer KV communication. KITE identifies a task-effective key layer using a receiver trajectory distortion criterion, transmits only the latent working memory associated with the key layer, and further uses the same layer as the entry point for autoregressive latent reasoning. Experiments on seven benchmarks across two model families and three model scales show that, compared with full-layer KV communication, KITE reduces communication volume by 28-36$\times$, achieves up to 3$\times$ end-to-end inference speedup, and improves accuracy by up to 23.3 percentage points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑