基于蒙特卡洛树搜索的多智能体系统自主修复
Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
AI总结:
该研究提出基于蒙特卡洛树搜索的MARS框架,将多智能体系统修复建模为搜索过程,结合分类学增强评估与诊断引导扩展,在StateMAS基准上性能优于现有方法,且令牌消耗相当。
AI中文摘要:
多智能体系统(MAS)正越来越多地被用于解决复杂任务。当输出不正确或不符合要求时,用户必须通过检查智能体轨迹(即故障归因)手动定位智能体错误,并提供反馈以优化输出(即修复)。尽管近期在MAS故障归因方面有一些研究,但从此类错误中恢复的自动化机制仍在很大程度上未被探索。为填补这一空白,我们提出了MARS,一种基于搜索的框架,它将MAS修复表述为蒙特卡洛树搜索(MCTS)过程,并通过结合分类学增强评估的诊断引导扩展,在庞大的潜在修复空间中进行导航。与通过完整展开评估完整模拟的标准MCTS不同,MARS通过部分展开评估智能体轨迹以减少令牌消耗。此外,我们引入了StateMAS,这是一个大规模MAS修复基准,包含1310条可重放的多智能体故障轨迹,涵盖四种智能体架构和四种大语言模型(LLM) backbone。在StateMAS上的实验表明,MARS在所有设置中始终优于最先进的方法,实现了3.0%至12.1%的绝对性能提升,同时保持了相当的令牌消耗成本。 ablation研究进一步证实,分类学增强评估和诊断引导扩展是实现这些性能提升的关键因素。
英文摘要:
Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0\% to 12.1\% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.