发表机构
Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对多智能体LLM系统的多错误归因问题,提出EDGE框架,通过构建错误依赖图并结合反事实回滚验证,在TRAIL、MAST等数据集上提升了归因性能,证明依赖结构是有效的诊断先验。
AI 中文摘要
大语言模型(LLM)智能体的故障通常包含多个相关错误,而非单一错误。现有归因方法通常仅识别负责的智能体、步骤或根本原因,但未显式建模错误间的依赖关系。我们提出EDGE,一种错误依赖图引导的多错误归因框架。EDGE从观测到的错误事件构建错误依赖图,并通过反事实回滚验证可靠的因果子集。该推理图引导两阶段LLM作为评判器的检测器进行错误归因,经干预验证的子图为解释和修复分析提供更可靠的基础。在TRAIL和MAST数据集上的实验表明,EDGE在大多数评估模型和设置中提升了类别级多错误归因性能;采用适配的Who&When风格提示的实验显示,该图在各类提示策略下均有帮助。这些结果表明,依赖结构是智能体故障诊断的有用先验,超出了孤立根本原因预测的范畴。
英文摘要
Large language model (LLM) agent failures often contain multiple related errors rather than a single mistake. Existing attribution methods usually identify a responsible agent, step, or root cause, but do not explicitly model dependency between errors. We introduce EDGE, an Error Dependency Graph-guided multi-Error attribution framework. EDGE constructs an error dependency graph from observed error events and validates a reliable causal subset through counterfactual rollout. The inference graph guides a two-stage LLM-as-judge detector for error attribution, and the intervention-validated subgraph provides a more reliable basis for explanation and repair analysis. Experiments on TRAIL and MAST show that EDGE improves category-level multi-error attribution across most evaluated models and settings. Experiments with adapted Who&When-style prompts show that the graph helps across prompting strategies. These results suggest that dependency structure is a useful diagnostic prior for agent failures beyond isolated root-cause prediction.
Journal refEMNLP 2026