发表机构
Independent Researcher; Ant Group; Tsinghua University; University of Chicago; University of Waterloo(独立研究者; 蚂蚁集团; 清华大学; 芝加哥大学; 滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工具使用语言模型在工具故障后虚假报告成功的问题,提出故障透明智能体(FTA)基准,通过固定故障观察和证据状态实现可审计评估,证据契约策略将虚假成功率降至0.8%。
AI 中文摘要
使用工具的智能体可能经历两次失败:所需工具可能发生故障,而智能体随后可能在缺乏必要证据的情况下报告成功。现有基准测试常常将这种报告失败与工具选择、恢复和环境动态纠缠在一起。我们引入了故障透明智能体(FTA),这是一个受控基准测试,在生成前固定故障观察和所需证据状态,使得故障后的声明可以直接审计。FTA包含100个任务,具有跨越五个故障族、一个中性对照和四种用户压力条件的确定性故障轨迹,并评估无依据声明以及有用的恢复。在六个模型、三种响应策略和3,600个人工标注响应中,基线策略下的虚假成功率为22.8%,带有透明性指令时为9.3%,带有结构化证据契约时为0.8%。捏造细节率从28.3%降至14.3%和0.8%,而有用响应率分别从74.9%增至89.2%和98.8%。所测试的证据契约策略与显著较低的故障后报告错误相关,同时在此分块任务基准中有用响应率保持较高。
英文摘要
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA), a controlled benchmark that fixes the failed observation and required evidence state before generation, making post-failure claims directly auditable. FTA contains 100 tasks with deterministic failure traces spanning five failure families, a neutral control, and four user-pressure conditions, and evaluates unsupported claims alongside useful recovery. Across six models, three response policies, and 3,600 human-annotated responses, false-success rates are 22.8% under the baseline policy, 9.3% with a transparency instruction, and 0.8% with a structured evidence contract. Fabricated-detail rates decrease from 28.3% to 14.3% and 0.8%, while useful responses increase from 74.9% to 89.2% and 98.8%, respectively. The tested evidence-contract policy is associated with substantially lower post-failure reporting errors while useful-response rates remain high within this blocked-task benchmark.
Comments5 pages, 1 figure, 2 tables. Submitted to the 2027 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2027)