发表机构
Qwen DianJin Team, Alibaba Cloud Computing; Beihang University(阿里云通义千问点睛团队; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DeFA提出依赖引导的失败归因框架,通过事件依赖图与失败传播图定位LLM智能体执行中的决定性错误,在多个基准上取得最优性能,并利用诊断反馈提升下游任务准确率6-15个百分点。
AI 中文摘要
LLM智能体执行中的错误与其可见后果之间可能相隔许多步骤,这使得决定性错误的定位需要同时理解步骤内容与步骤依赖关系。我们提出DeFA,一个依赖引导的智能体失败归因框架。DeFA首先将协议关系与语义依赖结合,构建覆盖轨迹的事件依赖图。随后,它识别可能违反任务要求的事件,并追踪其来源与后续影响,以构建失败传播图。最后,DeFA利用步骤证据及步骤在失败传播中的角色,识别决定性错误、责任智能体及错误类别。为支持长轨迹,DeFA将执行划分为多个片段,并将当前片段的详细内容与其他片段的摘要相结合,使局部诊断能够访问全局执行上下文。在Who and When及Who and When Pro文本子集上,DeFA在所有评估骨干模型上取得了最高的责任智能体与精确步骤准确率,并在Pro上取得了分类对齐方法中最高的失败模式准确率。进一步的图像与视频轨迹实验证明了其在多模态失败归因中的适用性。消融实验支持了分段、事件依赖图及失败传播图的贡献。将DeFA的诊断反馈用于Trace2Skill中的技能进化,在下游任务准确率上比原生管道提高了6-15个百分点,表明诊断还能支持智能体在后续任务上的改进。
英文摘要
Errors in LLM agent executions and their visible consequences can be separated by many steps, making decisive-error localization a matter of understanding both step content and step dependencies. We introduce DeFA, a dependency-guided framework for agent failure attribution. DeFA first combines protocol relations and semantic dependencies into an event dependency graph spanning the trajectory. It then identifies events that may violate task requirements and traces their sources and subsequent effects to construct a failure propagation graph. Finally, DeFA uses step evidence and the steps' roles in failure propagation to identify the decisive error, responsible agent, and error category. To support long trajectories, DeFA partitions executions into segments and combines the current segment's detailed content with summaries of the other segments, giving local diagnosis access to global execution context. Across Who and When and the Who and When Pro text subset, DeFA achieves the highest responsible-agent and exact step accuracy with all evaluated backbones, and the highest failure-mode accuracy among taxonomy-aligned methods on Pro. Further experiments on image and video trajectories demonstrate its applicability to multimodal failure attribution. Ablations support the contributions of segmentation, the event dependency graph, and the failure propagation graph. Using DeFA's diagnostic feedback for skill evolution in Trace2Skill improves downstream task accuracy by 6-15 percentage points over the native pipeline, showing that the diagnoses can also support agent improvement on subsequent tasks.
CommentsDeFA: Dependency-Guided Failure Attribution for LLM Agents