HELENA:面向多智能体系统的互补拓扑联合上的分层稀疏协调框架
HELENA:Hierarchical Sparse Coordination over a Union of Complementary Topologies for MAS
AI总结:
HELENA 是一种 MAS 框架,通过互补拓扑联合与分层稀疏协调抑制冗余噪声,在八个基准测试上实现最优结果,平均提升 3.47%,MMLU-Pro 最高提升 10.34%,难基准提升更显著且附加成本合理。
AI中文摘要:
基于大语言模型(LLM)的多智能体系统(MAS)通常仅优化单一拓扑结构,这将推理限制在狭窄轨迹内,限制了综合分析能力。若将多个拓扑简单合并为复合图,会引入无关连接间的冗余噪声传播,降低解决方案质量。为解决该困境,本文提出面向多智能体系统的互补拓扑联合上的分层稀疏协调框架(HELENA),这一多智能体框架通过依赖任务的稀疏执行平衡多样推理路径。HELENA 基于蒙特卡洛树搜索与确定性点过程筛选出互补候选拓扑,构建多智能体系统的联合图,拓宽推理轨迹以综合分析复杂问题;随后,分层稀疏协调模块在每一步仅激活一个稀疏子图,同时智能体交换压缩后的潜在摘要以抑制冗余噪声传播;最后,局部自精修阶段识别存在差异证据的决策单元,仅当对比证据同时确认可靠的解决方案侧存在失败且挑战者侧存在改进时,才对该单元进行重写。在八个基准测试上开展的实验显示,HELENA 在所有基准测试上均取得了最优结果,相比最强基线平均提升 3.47%,在 MMLU-Pro 上最高提升 10.34%,且在更难的基准测试上实现了更大提升,同时附加成本处于合理范围。
英文摘要:
LLM-based multi-agent systems (MAS) typically optimize a single topology, restricting reasoning to a narrow trajectory and limiting comprehensive analytical capacity. Naively merging multiple topologies into a composite graph introduces redundant noise propagation across irrelevant connections, degrading solution quality. To address this dilemma, we propose \textbf{Hierarchical Sparse Coordination over a Union of Complementary Topologies for MAS (HELENA)}, a multi-agent framework that balances diverse reasoning paths with sparse task-dependent execution. \helena{} constructs a union MAS graph from complementary candidate topologies selected via Monte Carlo Tree Search and Determinantal Point Process, broadening the reasoning trajectory for comprehensive analysis of complex problems. A Hierarchical Sparse Coordination module then activates only a sparse subgraph at each step while agents exchange compressed latent briefs to suppress redundant noise propagation. Finally, a Local Self-Refinement stage identifies decision units with discrepancy evidence and rewrites them only when contrastive evidence simultaneously confirms a reliable solution-side failure and a challenger-side improvement. Experiments across eight benchmarks show that \helena{} achieves state-of-the-art results on all benchmarks, with an average gain of \pctup{3.47} over the strongest baseline and up to \pctup{10.34} on MMLU-Pro, achieving larger improvements on harder benchmarks at a reasonable additional cost.