arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13463cs.AIcs.HCcs.LGcs.SE

根因归因是一个搜索问题:对长时程智能体失败的持续搜索

Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

  • Scale AI

机构由 AI 辅助整理,请以论文原文为准。

Harsh Raj, David Lee, Anas Mahmoud, Renxiong Wang, Razvan-Gabriel Dumitru, Chenguang Wang, Tong Zhao, Yunzhong He, Darvin Yi, Vipul Gupta

AI总结:

针对长时程AI智能体失败诊断中信息稀疏、分布广导致根因归因困难的问题,提出持续搜索迭代框架,在多个基准上显著提升归因性能,并证明有效搜索优于模型规模。

AI中文摘要:

随着AI智能体在长时程任务中的部署日益增多,产生了海量的执行日志。诊断这些记录中的失败对于可靠性至关重要,因为它将结果层面的信号转化为可操作的干预措施。数据的庞大规模使得人工审查不切实际,从而推动了对自动化根因归因(RCA)的需求。然而,使用LLM的自动化RCA方法诊断准确性较低,尤其是在执行轨迹变得更大时。这些方法之所以困难,是因为相关信息往往稀疏、分布在遥远的动作之间,并且与可见的失败脱节,从而将根因归因简化为一个大规模的搜索问题。现有的RCA方法通常依赖一次性LLM判断来从执行轨迹中诊断失败。虽然对于较短的轨迹有效,但这些判断者往往过早地确定一个看似合理的诊断,导致较长轨迹中的关键证据未被检查。我们引入了持续搜索(Continual Search),这是一个迭代框架,通过连续多轮推动判断者继续搜索未解决的诊断证据。我们在四个现有的RCA基准上评估了持续搜索。鉴于当前基准缺乏大规模执行轨迹,我们引入了MegaRCA-Mix来评估大规模RCA。MegaRCA-Mix提供了一个具有挑战性的测试平台,包含50个人工标注的失败试验,涵盖长时程、执行密集的任务。在多个基准套件和模型家族中,持续搜索持续提高了归因性能。例如,在MegaRCA-Mix上,它将GPT-5.5的F1分数提高了超过40%,从0.349提升到0.498。更有趣的是,在同一模型家族内,较低级别的模型甚至可以超越其较高级别的对应模型,这表明有效的搜索胜过了原始模型规模。

英文摘要:

The increasing deployment of AI agents in long-horizon tasks yields massive execution logs. Diagnosing failures within these records is crucial for reliability, as it transforms outcome-level signals into actionable interventions. The sheer scale of the data renders human review impractical, driving the need for automated root-cause attribution (RCA). However, automated RCA methods using LLMs suffer from low diagnostic accuracy, especially as execution traces grow larger. They struggle because relevant information is often sparse, distributed across distant actions, and disconnected from the visible failure, reducing root-cause attribution to a massive search problem. Existing RCA methods typically rely on one-shot LLM judgments to diagnose failures from execution traces. While effective for shorter trajectories, these judges tend to settle on a plausible diagnosis early, leaving critical evidence in longer traces unexamined. We introduce Continual Search, an iterative framework that nudges the judge, over successive turns, to keep searching for unresolved diagnostic evidence. We evaluate Continual Search across four existing RCA benchmarks. Recognizing the lack of massive execution traces in current benchmarks, we introduce MegaRCA-Mix to evaluate RCA at scale. MegaRCA-Mix provides a challenging testbed of 50 human-annotated failure trials spanning long-horizon, execution-heavy tasks. Across multiple benchmark suites and model families, Continual Search consistently improves attribution performance. On MegaRCA-Mix, for example, it improves Opus-4.8's F1 score by 29\%, from $0.471$ to $0.608$. More interestingly, within the same model family, lower-tier models can even surpass their higher-tier counterparts, demonstrating that effective search supersedes raw model scale.

↑