arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

路由、通信与推理:用于高效多智能体推理的门控路由与自适应深度

Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning

Sudipto Ghosh, Tanmoy Chakraborty

arXiv 2607.10836首次发表:更新:

发表机构

Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi; Department of Electrical Engineering, Indian Institute of Technology Delhi(印度理工学院德里分校雅迪人工智能学院; 印度理工学院德里分校电气工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对多智能体推理中未解决的问题,提出GRADE分层多智能体系统,用四个轻量级学习门控及CoGRPO训练方法,智能体模型可热插拔。该系统在多个任务上优于基线,消融实验明确关键因素,证明校准对热插拔必要。

AI 中文摘要

多智能体集成增加了有效参数和推理成本,却未解决三个基本问题:咨询哪些智能体、查询应在智能体层次结构中遍历多深、智能体间通信何时值得。我们提出了GRADE(用于高效推理的门控路由与自适应深度),这是一种分层多智能体系统,四个轻量级学习门共同控制智能体选择、层次深度、智能体间通信和分支修剪。训练使用CoGRPO(协作组相对策略优化),一种无评论家的方法,使GRPO适用于多智能体层次结构,并为参与一次展开的每个门和智能体分配共享优势信号。智能体模型来自可热插拔的专家注册表;每个智能体的校准图允许在推理时更换专家而无需重新训练。在平均约17B有效参数时,GRADE在GSM8K、MMLUPro和GPQA上优于所有基线,在MMLUPro上比最强基线高出4.8分且计算量减半。在AIME - 2025上,GRADE也保持竞争力。消融实验表明层次结构和掩码交叉注意力对准确性贡献最大,且每个智能体的校准对安全热插拔很必要。

英文摘要

Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We present GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning), a hierarchical multi-agent system in which four lightweight learned gates jointly govern agent selection, hierarchy depth, inter-agent communication, and branch pruning. Training uses CoGRPO (Collaborative Group-Relative Policy Optimization), a novel critic-free recipe that adapts GRPO to multi-agent hierarchies and assigns a shared advantage signal to every gate and agent that participated in a rollout. Agent models are drawn from a hot-swappable Expert Registry; per-agent calibration maps allow experts to be replaced at inference time without retraining. At $\sim$17B average active parameters, GRADE outperforms all baselines on GSM8K, MMLUPro, and GPQA, surpassing the strongest baseline by 4.8 points on MMLUPro at half the active compute. On AIME-2025, where model depth dominates, GRADE remains competitive to existing frameworks. Ablations isolate the hierarchy and masked cross-attention as the largest contributors to accuracy, and show that per-agent calibration is necessary for safe hot-swapping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑