发表机构
Mohamed bin Zayed University of Artificial Intelligence; University of Electronic Science and Technology of China; Nanjing University; Griffith University; University of New South Wales(穆罕默德·本·扎耶德人工智能大学; 电子科技大学; 南京大学; 格里菲斯大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种将智能体提示词转换为最小对比对的方法,发现工具调用决策由因果必要且充分的向量μ_Δ介导,分析动词通过抑制工具调用先验来阻止工具使用,该机制在多个模型家族中普遍存在。
AI 中文摘要
工具调用,即按需调用外部工具,是智能体大语言模型的核心能力,然而决定模型是调用工具还是直接响应的机制仍未被充分理解。智能体提示词冗长且高度结构化,将角色指令、工具模式、格式模板和用户请求组合成数百个词元,形成嘈杂且高度纠缠的上下文,其中没有单一可控变量可供机制分析。为获得这样的变量,我们提出一种方法,将复杂的智能体提示词转换为最小对比对,其中单个请求动词决定工具调用决策:将执行动词(如“写”)替换为分析动词(如“讨论”)可靠地翻转决策,表明该决策由紧凑的内部状态介导。我们在Python、Java和C++中构建了500个这样的配对提示词(300个用于机制分析,200个留作评估)。我们将该决策追溯到向量μ_Δ,该向量在因果上既必要又充分,并且能泛化到发现提示词之外的原生多轮τ^2-Bench轨迹和无动词请求。行为消融表明,脚手架建立了工具调用先验;Transcoder分解随后揭示,分析动词通过指示工具使用不必要的特征来抑制该先验,而执行动词则基本保持该先验不变。下游的脚手架读取注意力头和MLP特征读取结果状态,相同的机制在Qwen、Mistral和Granite家族的七个模型中重复出现。我们的代码可在https URL获取。
英文摘要
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textit{write}) with an analysis-verb (e.g., \textit{discuss}) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, $μ_Δ$, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn $τ^2$-Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at https://github.com/XijieGo/MI4ToolCalling.
CommentsAccepted by NeurIPS 2026 Main Poster