通过表示调控实现LLM智能体中可调节的工具调用率
Tunable Tool-Call Rates in LLM Agents via Representation Steering
浏览论文内容
中文总结 AI 辅助
本研究提出通过调控LLM残差流的线性方向,可在推理时无需额外训练调节工具调用率,使开放域问答准确率近翻倍,且能泛化到多种模型与工具。
中文摘要 AI 辅助
决定是否调用工具是大语言模型(LLM)智能体的核心能力,且该判断失误的代价高昂:不必要的调用会增加延迟、产生成本,还可能触发不可逆的副作用;而遗漏调用则会让模型在仅能通过工具调用才能回答的问题上给出自信的错误答案。现有方法如后训练和提示工程成本高昂,且难以在推理时修改。本文表明,指令微调模型是否调用工具可通过其残差流中的单个线性方向控制,该方向无需任何训练即可从模型自身的工具使用偏好信号中提取,并转化为无需修改提示的推理时干预措施。以强度α添加该方向可使调用率从接近0%单调提升至超过90%,同时保持调用格式正确。该调控可双向作用:降低强度会抑制调用,提高强度则会诱导新的调用,且这些调用恰好针对模型无法从自身知识中回答的问题。本文还表明,该方向可泛化到未见过的工具,其强度与各工具自身的方向相当,且不会偏向任何特定工具选择。通过实时工具执行,一次调控扫描即可绘制出成本/准确率的帕累托前沿,使开放域问答准确率几乎翻倍(从0.29提升至0.56);该方法可迁移到包含稠密模型、混合专家(MoE)模型和多模态架构的多种不同模型,无需任何训练。本文代码可在该https URL获取。
英文摘要
Deciding whether to call a tool is a core competence of an LLM agent, and a costly one to get wrong: needless calls add latency, accrue cost, and may trigger irreversible side effects, while missing calls leave the model confidently wrong on questions it could only answer through tool-calls. Models manage this balance poorly, both over-using and under-using tools. Existing methods such as post-training and prompt engineering are expensive and difficult to modify at inference time. We show that whether an instruction-tuned model calls a tool can be controlled by a single linear direction in its residual stream, extracted without any training from the model's own tool-use preference signal and turned into an inference-time intervention with no prompt change. Adding the direction with strength $α$ moves the call rate monotonically from near $0\% $ to over $90\%$ while keeping calls well-formed. The steering works in both directions: dialing it down suppresses calls, and dialing it up induces new calls that land precisely on the questions the model cannot answer from its own knowledge. We also show that the direction generalizes to unseen tools with strength comparable to each tool's own direction and without favoring any specific tool choice. With live tool execution, a single sweep of the steering traces a cost/accuracy Pareto frontier and nearly doubles open-domain QA accuracy ($0.29 \! \rightarrow \! 0.56$); the same recipe transfers across a diverse range of models spanning dense, MoE, and multimodal architectures, without any training. Our code is publicly available at https://github.com/YuqiChen4188/Steering-Tool-Use-Propensity.
发表机构
- UC Santa Cruz(圣克鲁兹加利福尼亚大学)
- UC Berkeley(加利福尼亚大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。