arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从证据到行动:工具使用型智能体为何失败

From Evidence to Action: How Tool-Using Agents Fail

Hongzhan Lin, Shidong Cao, Ziyang Luo, Wenhao Chai, Mong-Li Lee, Wynne Hsu

arXiv 2610.07753首次发表:更新:

发表机构

Princeton University; National University of Singapore; Hong Kong Baptist University; Amazon Web Services(普林斯顿大学; 新加坡国立大学; 香港浸会大学; 亚马逊云服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究工具使用型智能体从证据到行动的链条断裂问题,提出SafeActBench基准和证据账本,发现失败源于证据使用方式而非仅信息缺失。

AI 中文摘要

工具使用型智能体对外部状态做出重大改变,然而正确的结果并不能保证其行动事先得到了已确立证据的支持。我们研究了当智能体从决定是否行动转向执行单一行动和依赖工作流时,这一证据到行动链条在何处断裂。在十种模型-工具配置中,强大的静态行动评估能力可以与明显较弱的交互式执行能力并存。失败往往在执行之前就已开始:智能体在调查不完整时便停止,或在所需证据确立之前就采取行动。一旦获得所需证据,单一行动的执行通常是可靠的,而多行动工作流则会额外暴露未解决的前置条件和执行不完整的问题。为此分析,我们引入了SafeActBench,涵盖六个操作领域的656个案例和五种协议,这些协议从静态行动判断和调查后的不行动,逐步推进到单一行动和多行动工作流。一个基于来源的证据账本和确定性轨迹评估器追踪了哪些信息已确立、行动何时发生,以及下游依赖是否得到满足。这些结果表明,失败不仅源于信息缺失,还源于智能体在决定和执行行动时如何使用已确立的证据。

英文摘要

Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand. We study where this evidence-to-action chain breaks as agents move from deciding whether to act to executing single actions and dependent workflows. Across ten model-harness configurations, strong static action assessment can coexist with much weaker interactive execution. Failures often begin before execution: agents stop with incomplete investigation or act before required evidence is established. Once required evidence is obtained, single-action execution is usually reliable, while multi-action workflows additionally expose unresolved prerequisites and incomplete execution. For this analysis, we introduce SafeActBench, comprising 656 cases across six operational domains and five protocols that progress from static action judgment and investigated non-action to single- and multi-action workflows. A provenance-bound Evidence Ledger and deterministic trajectory evaluator track what information was established, when actions occurred, and whether downstream dependencies were satisfied. These results show that failures arise not only from missing information, but also from how agents use established evidence when deciding and executing actions.

Comments36 pages. Project page: https://safeact.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑