arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15242cs.AIcs.SE

LongRCA Bench:诊断长 horizon 智能体失败中的责任角色与根本原因

LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories

  • Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心)
  • Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)
  • Chongqing University(重庆大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • Singapore Management University(新加坡管理大学)
  • Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong… 展开作者

Yunfei Zhang, Boyu Feng, Changhua Pei, Zexin Wang, Zhihuang Peng, Xinlong Liu, Hengyue Jiang, Difeng Ma, Jiayi Zhang, Yongzhou Yao, Yanan Zhao, Fei Sun, Yintong Huo, Zhaoyang Liu, Jingjing Li, Gaogang Xie, Dan Pei

AI总结:

LongRCA Bench 是含1140条失败轨迹的长 horizon 智能体失败诊断基准,本文提出无需训练的 RCTA 方法,在责任角色归因和根本步骤定位上优于现有基线,表明需将二者作为独立评估目标。

AI中文摘要:

当长 horizon 智能体执行失败时,结果级评估仅会显示失败的结果,却不会揭示轨迹中决定性错误从何处进入。开发者随后必须检查完整执行过程,以识别责任角色并定位最早的决定性根本原因步骤。现有的失败归因基准大多聚焦于较短的轨迹,对数百个记录步骤的诊断仍未得到充分探索。我们推出 LongRCA Bench,它包含五个领域的1140条失败轨迹,且未注入错误。该基准提供了对责任角色和最早决定性根本原因步骤的独立评分的人工标签。轨迹的中位数包含145个步骤,最强基线仅达到13.2%的根本步骤精确准确率。我们进一步提出了根本原因轨迹归因(Root-Cause Trajectory Attribution,RCTA),这是一种无需训练的方法,它从片段摘要中检索候选错误步骤,并将其追溯到可用的更早交接指令。使用相同的主干、基准实例和评分协议,RCTA达到了51.1%的责任角色准确率和24.1%的根本步骤精确准确率。这些结果凸显了在长轨迹失败诊断中,需将责任角色归因和根本步骤精确定位作为独立目标进行评估。

英文摘要:

In long agent executions, an early error can persist through later actions and checks, while evidence needed to trace its origin is dispersed across the history. Short histories offer limited tests of recovering error origins across substantial subsequent execution. We introduce LongRCA Bench: 1,140 complete failed trajectories from five sources, all human-annotated for responsible roles and earliest decisive root-cause steps. Reference roots precede completion by a median of 48 recorded steps; 28.4% of trajectories contain over 100 subsequent steps. With DeepSeek-V4-Flash on the full benchmark, the strongest of five evaluated baselines achieves 13.2% exact root-step accuracy. We propose Root-Cause Trajectory Attribution (RCTA), a training-free method that organizes original candidate records and explicit handoff instructions for attribution. Segment summaries and a trajectory outline guide candidate retrieval; available handoff records supply upstream instruction context for the final instruction-execution comparison. With the same backbone and scoring protocol, RCTA reaches 24.1% exact root-step accuracy and 51.1% responsible-role accuracy. LongRCA-Mini provides 200 fixed trajectories for lower-cost comparative screening. Even with RCTA, fewer than one quarter of reference roots are recovered exactly.

补充信息

↑