理解并缓解使用工具的LLM智能体中的幻觉逃逸
Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究揭示工具使用LLM智能体中的幻觉逃逸现象,即现有缓解方法在调优配置上降幻觉但在其他配置上增幻觉,并提出EscapeGuard方法,通过冲突感知门控和注意力增强,在六个基准上将工具选择幻觉降低9.0个百分点,净改进达89.1%。
中文摘要 AI 辅助
大型语言模型(LLMs)日益作为调用外部工具的自主智能体。然而,这种能力引入了工具幻觉,即选择错误的工具或生成无效的调用。现有的缓解方法报告了显著的改进,但我们识别出一种先前被忽视的失败模式,我们称之为“幻觉逃逸”。这些方法在它们所调优的工具配置上减少了幻觉,但在其他配置上却增加了幻觉,从而抵消了收益。我们进一步研究这一现象,发现当模型固有的工具使用倾向与当前工具配置冲突时,幻觉会急剧上升,而现有方法强化而非抑制这些倾向,这反过来又导致了幻觉逃逸。基于这些发现,我们提出了EscapeGuard,一种无需训练的推理时方法,结合了冲突感知门控与配置派生的注意力增强,以缓解工具幻觉同时防止幻觉逃逸。在六个基准测试上,针对各种模型,EscapeGuard将工具选择幻觉降低了9.0个百分点,并抑制了幻觉逃逸,将跨配置均值降低了23.7个百分点,在配对查询评估中实现了89.1%的净改进。我们希望这项工作能鼓励超越单一工具配置的评估,并为更可靠的使用工具的LLM智能体铺平道路。
英文摘要
Large language models (LLMs) increasingly serve as autonomous agents that invoke external tools. However, this capability introduces tool hallucination, selecting incorrect tools or generating invalid calls. Existing mitigation methods report substantial improvements, yet we identify a previously overlooked failure mode that we term Hallucination Escape. These methods reduce hallucination on the tool configuration they are tuned on but increase it on other configurations, canceling out the gain. We further investigate this phenomenon and find that hallucination rises sharply when a model's intrinsic tool-use tendencies conflict with the current tool configuration, and that existing methods reinforce rather than suppress these tendencies, which in turn contributes to hallucination escape. Building on these findings, we propose EscapeGuard, a training-free inference-time method that combines conflict-aware gating with configuration-derived attention enhancement to mitigate tool hallucination while preventing hallucination escape. Across six benchmarks on various models, EscapeGuard reduces tool-selection hallucination by 9.0 pp and suppresses hallucination escape, lowering the cross-configuration mean by 23.7 pp and achieving an 89.1% net improvement in paired-query evaluation. We hope this work can encourage evaluation beyond a single tool configuration and pave the way for more reliable tool-using LLM agents.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。