arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BANDMAS:面向带宽高效多智能体协作的因果启发式语义分组调度

BANDMAS: Causality-Inspired Semantic Packet Scheduling for Bandwidth-Efficient Multi-Agent Collaboration

Jiangwen Dong, Wanyu Lin

arXiv 2608.00458首次发表:更新:

AI 中文总结

该研究针对多智能体协作中通信流量大、开销高的问题,提出BANDMAS框架,通过因果启发式语义分组调度减少应用层流量,在三类数据集上实现了显著的流量降低并取得最优任务指标。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统通过各专用智能体间交换消息聚合信息来做出决策。将每个生成的消息在智能体间转发会增加应用层流量,同时为智能体处理带来大量输入令牌,可能提升推理延迟与计算开销。现有方法尝试通过修剪智能体或丢弃冗余消息解决上述问题,但这类智能体级或消息级优化会导致最终决策缺乏足够证据,或仍存在冗余消息传输。为应对这些挑战,我们提出BANDMAS,这是一个将智能体间通信建模为面向任务的流量的多智能体协作框架,可通过因果启发式重放估值实现高效传输。具体而言,我们通过分析消息的语义特征(如证据与请求)将其分解为多个数据包,系统仅在数据包预测的重放衍生贡献超过其资源成本时才传输它们。因此,BANDMAS能够在遵守带宽、延迟、截止日期及接收方上下文约束的同时自适应调度通信数据包。在冻结的Qwen3-4B模型于SciFact、HotpotQA和FanOutQA数据集上的流量实验中,我们的框架在选定的上限下将应用层字节减少了53.2%至77.3%,并在所有三个工作负载的受约束方法中获得了最高的平均任务指标。

英文摘要

LLM-based multi-agent systems make decisions based on the aggregated information via exchanging messages across specialized agents. Forwarding every generated message among agents increases application-layer traffic. Yet, it introduces tremendous input tokens for agent processing, potentially raising inference latency and computational overhead. Existing approaches attempt to address the above issues by pruning agents or discarding redundant messages. Nevertheless, such agent-level or message-level optimization results in insufficient evidence supporting for final decisions or still containing redundant message transmissions. To address these challenges, we propose BANDMAS, a multi-agent collaboration framework that models inter-agent communications as task-oriented traffic, which enables efficient transmission via causality-inspired replay valuation. Specifically, we decompose messages into several data packets by analyzing their semantic features such as evidence and requests. The system only transmits these packets if their predicted replay-derived contribution exceeds their resource cost. Consequently, BANDMAS is able to adaptively schedule communication packets while adhering to bandwidth, latency, deadline, and receiver context constraints. On frozen Qwen3-4B traffic across SciFact, HotpotQA, and FanOutQA, our framework reduces application-layer bytes by 53.2\% to 77.3\% at selected caps and attains the highest mean task metric among constrained methods on all three workloads.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑