arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19425cs.AIcs.CRcs.SE

LLM智能体中的工具幻觉的封闭世界消解

Closed-World Resolution Against Tool Hallucination in LLM Agents

Laxmipriya Ganesh Iyer

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对LLM智能体调用不存在工具的结构性盲点,提出封闭世界消解器与五类幻觉分类法,并通过基准测量证明幻觉防御必须先于门控。

中文摘要 AI 辅助

工具增强的大型语言模型(LLM)智能体以一种任何工具选择或工具安全方法都无法解决的方式失败:它们调用不存在的工具,并传递没有任何模式声明的参数。现有的防御措施要么选择正确的工具(选择),要么约束智能体对真实工具可能采取的操作(门控),两者都预设了发出的调用指向一个真实工具。我们表明这是一个结构性盲点:幻觉调用在构造上不是任何门控所做的决策,因此没有门控可以拒绝它。本文主要是一项测量和基准研究。我们给出了工具幻觉的五类分类法(H1-H5),并作为参考点,提出了“消解阶梯”(Resolution Rung):一种无需训练、封闭世界的消解器(注册表成员资格加签名检查),其意义在于它必须处于的位置,而非其计算内容。我们证明了幻觉防御必须先于任何因果门控,并刻画了唯一不可约的残留(借用参数在模式上与有效调用不可区分)。在两种调用界面下,跨十个托管模型,我们测量到322个真实幻觉;伪造工具调用集中在无约束的原始JSON界面(34对3),且模型规模无济于事(一个675B模型与7-8B模型相当)。然后我们扩展到模型上下文协议(Model Context Protocol),其中将多个服务器合并到一个命名空间会创建单个注册表无法表达的幻觉表面(第二个分类法,M1-M5);在实时MCP界面上,我们测量到154个幻觉,包括来自在单注册表界面上表现干净的前沿模型,因为冲突和遮蔽是合并的结构性结果。我们发布了带版本号的幻觉工具基准(Hallucinated-Tools Benchmark, HTB),以便任何消解器在提交之间具有可比性。

英文摘要

Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all. We show this is a structural blind spot: a hallucinated call is by construction not a decision any gate made, so no gate can reject it. This paper is primarily a measurement and benchmark study. We give a five-class taxonomy of tool hallucination (H1-H5) and, as a reference point, the Resolution Rung: a training-free, closed-world resolver (registry membership plus a signature check) whose interest is where it must sit, not what it computes. We prove hallucination defense must precede any causal gate, and characterize the one irreducible residue (borrowed arguments schema-indistinguishable from a valid call). Across ten hosted models under two invocation surfaces we measure 322 genuine hallucinations; fabricated-tool calls concentrate on the unconstrained raw-JSON surface (34 vs. 3), and model scale does not help (a 675B model matches a 7-8B one). We then extend to the Model Context Protocol, where merging several servers into one namespace creates hallucination surfaces a single registry cannot express (a second taxonomy, M1-M5); on the live MCP surface we measure 154 hallucinations, including from frontier models that were clean on the single-registry surface, because collisions and shadowing are structural to the merge. We release the versioned Hallucinated-Tools Benchmark (HTB) so any resolver is comparable across submissions.

↑