arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37017cs.CL

LatCom:用于高效多智能体协作的跨智能体潜在压缩

LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration

Shinan Zhang, Tao Zhang, Qihui Zhu, Mengjie Zhang, Dong Jin, Yunpeng Hou, Shuangwu Chen, Xiaobin Tan, Quan Zheng, Jian Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多智能体潜在协作中跨智能体冗余问题,提出LatCom框架,将多个发送方潜在表示压缩为固定槽位,分两阶段训练,实现2.46倍加速并减少70.3%输出令牌。

中文摘要 AI 辅助

基于大语言模型的多智能体系统(MAS)越来越多地采用潜在协作方式,以避免自然语言通信中的信息损失和重复的编码-解码开销。然而,直接转发所有发送方的潜在表示会使接收方的上下文规模随智能体数量和推理长度同时增长,从而增加计算量、内存占用和协作延迟。一种自然的解决方案是潜在压缩。但我们发现,现有的潜在压缩方法中,跨智能体的冗余问题仍未解决——这些方法通常独立压缩每个发送方的潜在表示,然后将结果拼接起来。我们提出LatCom,一种用于高效多智能体潜在协作的跨智能体潜在压缩框架。LatCom将多个发送方的潜在表示映射到固定数量的、接收方可读且与任务相关的槽位中。它并非重建所有发送方的隐藏状态,而是针对接收方的任务效用优化压缩后的潜在表示。LatCom分两个阶段训练压缩器:单发送方可读性学习首先建立一个可被冻结接收方解释的潜在接口;多发送方融合学习随后训练压缩器融合互补证据并消除跨智能体的冗余。在使用Qwen3-4B的多个基准上的实验表明,与LatentMAS相比,LatCom平均实现了2.46倍的推理加速,并将输出令牌使用量减少了70.3%,同时保持了相当的平均准确率。

英文摘要

LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatenate the results. We propose LatCom, a cross-agent latent compression framework for efficient multi-agent latent collaboration. LatCom maps multiple sender latents into a fixed number of receiver-readable and task-relevant slots. Rather than reconstructing all sender hidden states, it optimizes the compressed latents for receiver-side task utility. LatCom trains the compressor in two stages: single-sender readability learning first establishes a latent interface interpretable by the frozen receiver, and multi-sender fusion learning then trains the compressor to fuse complementary evidence and remove redundancy across agents. Experiments on multiple benchmarks with Qwen3-4B show that LatCom achieves an average 2.46x inference speed-up over LatentMAS and reduces output token usage by 70.3% while maintaining comparable average accuracy.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑