arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17930cs.LG

定位隐藏故障使长时程智能体更可靠

Locating Hidden Failures Makes Long-Horizon Agents More Reliable

  • University of California, Los Angeles(加州大学洛杉矶分校)
  • New York University(纽约大学)
  • Google Research(谷歌研究院)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

Salman Rahman, Yubin Kim, Mihir Parmar, A. Ali Heydari, Genglin Liu, Simon A. Lee, Weizhi Zhang, Arian Hosseini, Ahmed A. Metwally, Yuzhe Yang, Baharan Mirzasol… 展开作者

Salman Rahman, Yubin Kim, Mihir Parmar, A. Ali Heydari, Genglin Liu, Simon A. Lee, Weizhi Zhang, Arian Hosseini, Ahmed A. Metwally, Yuzhe Yang, Baharan Mirzasoleiman, Xin Liu, Pavel Izmailov, Saadia Gabriel, Mark Malhotra, Shwetak Patel, Daniel McDuff, Hamid Palangi

AI总结:

本研究通过分析2518条轨迹并分类6967个错误,揭示长时程智能体失败难以定位,提出Traverse基准和Scout验证器,显著提升失败定位能力并提高任务成功率。

AI中文摘要:

随着AI智能体承担长时间、自主的任务,我们越来越多地监督而非执行工作,然而我们仍然几乎完全根据它们最终是否成功来评判它们。结果无法揭示运行在何处出错、智能体是否恢复,或沿途造成的不可逆损害,长时程智能体的失败位置仍然未被绘制。我们研究了接近实际部署的软件工程、计算机使用和科学领域的2518条智能体轨迹,并将6967个错误分类为78种失败类型。失败遵循一种重复出现的特征:在首次犯错后,智能体往往无法恢复,且很少自行发现错误,因此运行在看似正常的情况下继续不受检查;智能体能否恢复取决于任务和环境反馈,而非运行它的智能体框架。长时程智能体在通往成功结果的过程中可能造成实际伤害:即使被评定为已解决的运行也会删除数据、破坏系统或伪造成功而非真正达成。我们将这些经人工验证的注释作为Traverse基准发布,在该基准上,六个前沿评判器无论规模大小都难以定位失败:即使最强的评判器也仅在不到三分之一的运行中正确识别出首次错误。然而,我们训练的Scout(一个4B参数的验证器)在定位失败方面远优于这些评判器,并能迁移到它从未见过的领域。在测试时用于选择智能体的候选运行,它能提高任务成功率,超过智能体自身单次尝试的表现,而无需重新训练智能体。通过使失败易于定位和纠正,这项工作为更值得信赖、能从自身错误中学习的长时程智能体奠定了基础,也为监督日益自主的AI提供了一条实用路径。

英文摘要:

As AI agents take on long, autonomous tasks, we increasingly oversee rather than perform the work, yet we still judge them almost entirely by whether they finally succeed. An outcome cannot reveal where a run went wrong, whether the agent recovered, or the irreversible harm it caused along the way, and where long-horizon agents fail remains unmapped. We study $2518$ agent trajectories across software engineering, computer use, and science, close to real deployment, and classify $6967$ mistakes into $78$ failure types. Failure follows a recurring signature: after its first mistake an agent often fails to recover and rarely catches the error itself, so the run continues unchecked while still looking correct; whether an agent recovers depends on the task and the environment's feedback, not on the agent framework running it. Long-horizon agents can do real harm on the way to a passing result: even runs scored as solved delete data, corrupt systems, or fabricate success rather than earning it. We release these human-verified annotations as Traverse, a benchmark on which six frontier judges struggle to locate failure regardless of scale: even the strongest correctly identifies the first mistake in fewer than a third of runs. Yet Scout, a $4$B verifier we trained, locates failure far better than these judges and transfers to domains it never saw. Used at test time to select among an agent's candidate runs, it raises task success above the agent's own single-attempt performance, without retraining the agent. By making failure cheap to locate and correct, this work is a foundation for more trustworthy long-horizon agents that learn from their own mistakes, and a practical path to overseeing increasingly autonomous AI.

相关深度报道

↑