arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体过度信任工具:衡量对不可靠工具的依赖

Agents' Overreliance on Unreliable Tools

Hoyeol Yang, Woojung Song, Taewon Kim, Jonghyun Song, Seoyeon Park, Yohan Jo

arXiv 2609.05587首次发表:更新:

发表机构

Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过破坏三个工具的输出,评估十四个大语言模型,发现智能体对不可靠工具存在高度过度信任(采纳率最高达68%),并测试了三种干预措施但均未能一致缓解该问题。

AI 中文摘要

现有的工具使用智能体评估主要衡量智能体是否能够借助工具成功完成各种任务。这些评估通常假设工具返回的信息是可靠的。然而,现实系统中的工具返回结果可能看似合理却不正确。我们通过使用三个工具(网络搜索、LLM子智能体委派和代码执行)评估十四个大语言模型,研究智能体如何应对不可靠的工具返回结果。对于每个工具,我们破坏其返回结果,并衡量智能体是否在最终答案中采纳被破坏的内容。智能体在所有三种设置中都表现出高度的过度信任:每个工具的采纳率平均值超过三分之一,其中网络搜索的采纳率达到68.0%。对推理轨迹的分析揭示了一个尤为令人担忧的失败模式:智能体常常识别出冲突,甚至内部恢复出正确答案,却仅呈现被破坏的答案而不向用户发出警告。为了缓解智能体对工具返回结果的过度信任,我们在三个层面进行干预:用户提示、工具提供者的元数据以及智能体构建者的后训练。尽管某些干预措施对特定模型或工具有帮助,但没有任何一种干预能跨工具一致地缓解过度信任。这些发现将不可靠工具上的过度信任确认为一种严重且持久的失败模式,促使评估和干预措施使智能体能够验证工具输出并透明地传达未解决的冲突。

英文摘要

LLM agents use tools to access information and perform computations beyond their parametric knowledge. Existing tool-use benchmarks evaluate whether agents select and call the right tools, assuming that tool returns are reliable. However, tools can return plausible but incorrect outputs. We evaluate 14 models with three tools (web search, an LLM sub-agent, and a code executor), corrupting their returns to examine whether agents overrely on unreliable tools. Agents adopt corrupted returns at high rates, with mean adoption exceeding one third for every tool and reaching 68.0% for web search. Corrupted returns also frequently override correct answers that agents give without tools, and this reliance persists in more complex tasks and in tasks that combine multiple tools. Agents often make more tool calls under corrupted returns, and reasoning traces show them questioning a return and even stating the correct answer. Nevertheless, they still pass the corrupted content to users, rarely warning them of the conflict. We further test interventions spanning prompts, tool metadata, post-training, and activation steering. Verification prompts reduce adoption across all three tools while largely preserving accuracy with correct returns, and reliability labels provide partial mitigation in web search. The tested post-training methods offer limited gains, and the effect of activation steering depends on the model and intervention strength. Overall, our findings provide a basis for diagnosing tool overreliance and for developing agents that assess tool returns rather than assuming their reliability.

Comments49 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑