发表机构
College of Intelligence and Computing, Tianjin University; School of New Media and Communication, Tianjin University; College of Computer Science and Artificial Intelligence, Fudan University(天津大学智能与计算学部; 天津大学新媒体与传播学院; 复旦大学计算机科学技术与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于LLM的多智能体系统故障归因的浅层归因与上下文退化问题,提出无需训练的DCFA框架,结合全局因果依赖图与局部反事实推理,在Who&When基准上使步骤级准确率最高提升8.27%。
AI 中文摘要
近年来,基于大语言模型(LLM)的多智能体系统发展迅速。尽管这类系统具有应用前景,但仍较为脆弱,频繁出现推理与协调错误,进而引发系统级故障。此类系统的故障归因需追踪智能体间的自然语言交互,以识别决定性错误——即最早的、经修正后可逆转系统故障的动作。存在两大关键挑战:1)浅层归因:现有方法往往仅捕捉到次要偏差,如检索不完整或格式错误(这类错误可通过验证机制修正),却遗漏了系统故障的决定性原因;2)上下文退化:随着系统轨迹长度增加,模型推理能力迅速下降。为应对这些挑战,本文提出DCFA,一种无需训练的故障归因框架。DCFA整合了两个模块:全局模块,从系统轨迹构建结构化的因果启发依赖图,以识别初始决定性错误;局部模块,应用局部反事实启发式推理优化因果启发归因。在Who&When基准数据集上针对六种LLM开展的实验显示,DCFA的步骤级准确率较现有最优基线提升了至多8.27%。
英文摘要
Large language model (LLM)-based multi-agent systems have experienced rapid growth in recent years. Despite their promise, such systems remain fragile, frequently exhibiting reasoning and coordination errors that can lead to system-level failures. Failure attribution in such systems relies on tracing natural language interactions among agents to identify the decisive error, which refers to the earliest action whose correction can reverse system failure. There are two key challenges: 1) Shallow attribution: Existing methods often capture only minor deviations, such as incomplete retrievals or formatting errors, which verification mechanisms can correct, while missing the decisive cause of system failure. 2) Contextual degradation: As the length of the system traces increases, the model's reasoning ability rapidly deteriorates. To address these challenges, we propose DCFA, a training-free framework for failure attribution. DCFA integrates a global module that constructs structured causal-inspired dependency graphs from system traces to identify the initial decisive error, and a local module that applies local counterfactual-inspired reasoning to refine causal-inspired attribution. Experiments on the Who&When benchmark across six LLMs show that DCFA improves step-level accuracy by up to 8.27% over state-of-the-art baselines.