arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

修复还是重采样?重新思考LLM多智能体系统中的故障调试

Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems

Zhongwen Luan, Xiaoyu Zhang, Ming Hu, Yue Yang, Jiongchi Yu, Xiaohong Chen

arXiv 2608.25920首次发表:更新:

发表机构

East China Normal University; Nanyang Technological University; Singapore Management University; Xi’an University of Architecture and Technology(华东师范大学; 南洋理工大学; 新加坡管理大学; 西安建筑科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对LLM多智能体系统的故障调试,提出SymTrace评估框架与SymFail数据集,发现现有无指导重运行方法不可靠,提出的症状驱动干预方法可显著提升故障修复率。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统(MAS)正越来越多地应用于长程复杂任务,其可靠性已成为阻碍其实际部署的核心瓶颈。现有的MAS调试与修复方法通常依赖于重新运行和重采样整个执行轨迹。然而,一个根本问题仍待解答:这些方法是通过因果关系修复MAS故障,还是仅利用LLM采样的随机性进行随机修复?为评估MAS修复方法的有效性,我们引入SymTrace,这是一种受控评估框架,可记录MAS执行轨迹并建立干预锚点。在重放过程中,它利用记录的日志有效重建锚点之前的执行过程,仅重新生成下游轨迹,从而实现MAS故障的可靠复现。我们进一步构建了SymFail数据集,包含536个人类标注的故障轨迹,带有图关联的位置、类别和追踪证据。基于这些基础,我们在三个主流MAS框架上开展了大规模实证研究。我们的发现表明,现有的无指导重运行方法极不可靠,其故障复现率和修复率分别仅为67.97%和6.90%。基于这些发现,我们进一步探索了症状驱动干预方法的有效性,该方法成功修复了20.15%的故障案例,较现有最先进的修复方法提升了191.89%。本研究旨在为MAS调试与修复研究提供可操作的见解,为多智能体系统的稳健部署铺平道路。

英文摘要

As large language model (LLM)-based multi-agent systems (MASs) are increasingly applied to long-horizon complex tasks, their reliability has emerged as the core bottleneck hindering their real-world deployment. Existing MAS debugging and repair methods typically rely on rerunning and resampling the entire execution trajectory. However, a fundamental question remains to be answered: do these methods causally repair MAS failures or merely stochastically repair by leveraging the randomness of LLM sampling? To evaluate the effectiveness of MAS repair methods, we introduce SymTrace, a controlled evaluation framework that records the MAS execution trajectory and establishes intervention anchors. During replay, it effectively reconstructs the execution before the anchor using recorded logs and only regenerates the downstream trajectory, thereby enabling the reliable reproduction of MAS failures. We further construct the dataset SymFail, comprising 536 human-annotated failure trajectories with graph-linked locations, categories, and trace evidence. Based on these foundations, we conduct a large-scale empirical study across three mainstream MAS frameworks. Our findings reveal that existing unguided rerun methods are highly unreliable, exhibiting low failure reproduction and repair rates (only 67.97% and 6.90%, respectively). Building upon these findings, we further explore the effectiveness of a symptom-driven intervention method, which successfully repairs 20.15% of the failed cases (a 191.89% improvement to state-of-the-art repair methods). This study aims to provide actionable insights for MAS debugging and repair research, paving the way for the robust deployment of multi-agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑