arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于System One引导的计算分工的令牌高效多智能体协作

Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor

Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang

arXiv 2610.08155首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Southwestern University of Finance and Economics(电子科技大学; 西南财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出S1-MAS框架,通过System One引导的计算分工将协调决策与昂贵推理解耦,在七个基准上显著降低令牌消耗和延迟,同时保持高准确性。

AI 中文摘要

基于大型语言模型(LLM)的多智能体系统(MAS)已成为处理复杂信息检索和推理任务的一种有前景的范式,它通过专业智能体之间的协作解决问题。然而,现有的MAS框架将任务推理与协调操作(包括任务选择、角色分配、消息路由和上下文管理)紧密耦合。随着交互的增加,使用强大的LLM进行这些有界控制决策会带来大量的令牌开销和延迟,限制了智能体Web服务的可扩展性。在本文中,我们研究了协调是否可以在不牺牲协作性能的情况下与昂贵的推理解耦。我们提出了S1-MAS,一种基于System One引导的计算分工的令牌高效多智能体框架。S1-MAS将有界的协调决策分配给轻量级的System One模型,同时将开放式推理保留给有能力的LLM工作者。具体来说,一个轻量级控制器选择检查条件、选择后续任务并确定终止,而一个紧凑的阅读器从授权来源检索与条件相关的证据以支持这些决策。通过决策-证据循环,选定的任务动态地确定工作者角色和来源访问权限,从而实现无需任务特定训练的适应性协作。在七个不同基准上的大量实验表明,S1-MAS在显著降低推理成本的同时实现了优越的准确性。在与AgentVerse、DyLAN和SelfOrg在七个基准上的个体比较中,S1-MAS将GPT-4o的令牌消耗减少了44.9%-97.2%,并将测量的端到端延迟减少了37.8%-93.0%。这些结果突显了其在可扩展且成本效益高的智能体Web应用中的潜力。

英文摘要

Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑