arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的多智能体系统中故障归因的误差传播建模

Error-Propagation Modeling for Failure Attribution in LLM-Based Multi-Agent Systems

Jiaqi Liao, Yuanzhao Zhai, Huanxi Liu, Xu Zhang, Zheming Zhuang, Dawei Feng, Bo Ding, Huaimin Wang

arXiv 2610.11600首次发表:更新:

发表机构

College of Computer Science and Technology, National University of Defense Technology; School of Mechanical Engineering, Tianjin University(国防科技大学计算机学院; 天津大学机械工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对基于大语言模型的多智能体系统故障归因难题,提出EMFA方法,通过建模误差传播与循环,结合候选筛选和反事实验证,在Who&When基准上实现最优步骤级归因准确率并提升了相关指标。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统(MAS)正越来越多地被用于通过协同推理、工具使用及与外部资源交互来解决复杂任务。然而,此类系统中的故障归因仍具挑战性,因为观测到的结果往往无法直接揭示导致执行失败的错误。本研究的归因目标是决定性错误,定义为智能体-步骤对,纠正该对即可恢复失败的执行。现有方法大多仅识别可疑步骤,未明确建模错误如何在交互中传播或在未解决的循环中持续存在,导致难以区分决定性错误与下游故障症状。我们提出用于故障归因的误差传播建模(EMFA),EMFA构建失败轨迹的结构化表示,对级联传播和持续交互循环均进行建模,并采用感知传播的候选筛选及反事实验证来识别决定性智能体-步骤对。在Who&When基准上,EMFA实现了最先进的步骤级归因准确率,且在智能体层面仍具竞争力;其在手工构建子集和算法生成子集上的步骤级结果较此前最佳水平分别提升了3.45和4.40个百分点。

英文摘要

LLM-based multi-agent systems (MASs) are increasingly used to solve complex tasks through coordinated reasoning, tool use, and interaction with external resources. However, attributing failures in such systems remains challenging because the observed outcome often does not directly reveal the error responsible for the failed execution. In this work, the attribution target is the decisive error, defined as the agent--step pair whose correction would recover the failed execution. Existing approaches largely identify suspicious steps without explicitly modeling how errors propagate across interactions or persist in unresolved loops, making decisive errors difficult to distinguish from downstream failure symptoms. We propose \textbf{E}rror-Propagation \textbf{M}odeling for \textbf{F}ailure \textbf{A}ttribution (\textbf{EMFA}). EMFA constructs a structured representation of the failed trajectory, models both cascading propagation and persistent interaction loops, and uses propagation-aware candidate screening followed by counterfactual verification to identify the decisive agent--step pair. On the Who\&When benchmark, EMFA achieves state-of-the-art step-level attribution accuracy and remains competitive at the agent level. It improves the previous best step-level results by 3.45 and 4.40 percentage points on the Hand-Crafted and Algorithm-Generated subsets, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑