arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从异常到失败:构建用于智能体轨迹诊断的因果错误图

From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis

Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen, Jiayi Gui, Dayong Yang, Wenbo Yu, Haoke Zhang, Jie Tang, Cunxiang Wang

arXiv 2609.32514首次发表:更新:

AI 中文总结

针对智能体轨迹诊断中异常、错误与失败混淆及因果建模缺失的问题,提出CEG-Agent框架,通过构建因果错误图并配套CEG-Bench基准,实现高精度因果故障归因,达到最先进性能。

AI 中文摘要

由大语言模型驱动的智能体正越来越多地部署在复杂应用中,其中长程智能体轨迹使得故障难以诊断。现有的轨迹诊断方法常常混淆异常、错误和失败,导致诊断目标模糊;它们也缺乏对因果相关错误如何传播并放大为最终任务失败的结构化建模,从而导致不可靠的失败归因。为解决这些问题,我们提出了CEG-Agent,一个用于智能体轨迹因果诊断的工具增强型智能体框架。具体而言,CEG-Agent引入了异常、错误和失败的明确分类法,并构建了因果错误图(CEGs),这是一种统一的类型化表示,通过因果关系将执行事件、诊断节点和失败结果联系起来。为了评估因果轨迹诊断,我们进一步构建了CEG-Bench,这是一个完全由智能体标注的基准,其高置信度、共识衍生的CEG标注通过对抗性智能体裁决协议(AAAP)获得。我们将所得标注与专家策划的人工金标准进行验证,结果显示与自动标注高度一致。在CEG-Bench上的实验表明,CEG-Agent在语义宽松和结构精确两种评估标准下均达到了最先进的性能。我们的代码已公开。

英文摘要

LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally relevant errors propagate and amplify into final task failures, resulting in unreliable failure attribution. To address these problems, we propose CEG-Agent, a tool-augmented agentic framework for causal diagnosis of agentic traces. Specifically, CEG-Agent introduces an explicit taxonomy of anomalies, errors, and failures, and constructs Causal Error Graphs (CEGs), a unified typed representation that links execution events, diagnostic nodes, and failure outcomes through causal relations. To evaluate causal trace diagnosis, we further construct CEG-Bench, a fully agent-annotated benchmark with high-confidence, consensus-derived CEG annotations obtained through an Adversarial Agentic Adjudication Protocol (AAAP). We validate the resulting annotations against an expert-curated human gold set, which shows close agreement with the automatic annotations. Experiments on CEG-Bench demonstrate that CEG-Agent achieves state-of-the-art performance under both semantically relaxed and structurally exact evaluation criteria. Our code is publicly available.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑