发表机构
Renmin University of China; Ant Group(中国人民大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型驱动多智能体系统的高失败率问题,提出DoCtOR框架,通过自动化失败归因识别决定性错误智能体,仅让其针对性反思,在多个数据集上提升成功率且优于基线方法。
AI 中文摘要
由大语言模型驱动的多智能体系统(MAS)在复杂任务中展现出潜力,但失败率较高。当前MAS的自动反思方法要求所有智能体对失败进行反思,却忽略了一个关键现实:失败通常源于某个将任务引向错误方向的特定智能体,即决定性错误智能体,而其他智能体仅履行常规职责。强制要求行为正常的智能体进行反思会用错误见解污染它们的记忆。因此,我们提出DoCtOR(Diagnose-then-Correct PPO-enhanced Reflection,即诊断后修正的PPO增强型反思),这是一种增强多智能体协作的新型反思框架。DoCtOR首先通过自动化失败归因识别决定性错误步骤和决定性错误智能体,然后采用反事实推理生成修正后的决定性错误步骤,最后仅让决定性错误智能体生成针对性反思。实验结果显示,DoCtOR在HotPotQA、ChartQAPro和Mind2Web数据集上的初始成功率分别提升了22%、26%和27%,优于Reflexion、Retroformer和COPPER。我们进一步验证了诊断后修正范式的通用性,并证明在低资源场景下,将反思聚焦于决定性错误步骤之后的推理步骤,可达到对完整失败轨迹进行反思的可比质量。
英文摘要
Multi-agent systems (MAS) powered by large language models have shown promise for complex tasks but suffer from high failure rates. Current self-reflection methods for MAS require all agents to reflect upon failure, overlooking a critical reality: failures typically stem from a specific agent leading the task astray, namely the decisive error agent, while others merely fulfill their regular duties. Forcing regular-behaving agents to reflect contaminates their memory with wrong insights. Hence, we propose DoCtOR (Diagnose-then-Correct PPO-enhanced Reflection), a novel reflection framework that enhances multi-agent collaboration. DoCtOR first identifies the decisive error step and decisive error agent through automated failure attribution, then employs counterfactual reasoning to generate a corrected decisive error step, and finally engages only the decisive error agent to produce targeted reflections. Experimental results show DoCtOR achieves 22%, 26%, and 27% improvements over initial success rates on HotPotQA, ChartQAPro, and Mind2Web datasets, outperforming Reflexion, Retroformer, and COPPER. We further establish the generalizability of our diagnose-then-correct paradigm and demonstrate that in low-resource settings, focusing reflection on reasoning steps after the decisive error step achieves comparable quality to reflecting on the complete failure trajectory.
CommentsAccepted by EMNLP 2026 main