arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16268cs.CL

虚假工具使用:强化学习代理何时学到错误的行动理由

Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act

Yiwei Yang, Haoxiang Zhang, Bingbing Wen, Yao Lu, Yuchen Wu, Lei Zhang, Julian McAuley, Pan Lu, Bill Howe

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨强化学习代理基于表面线索而非真实需求调用工具的虚假工具使用问题,提出工具必要性奖励机制,有效抑制捷径行为并保持任务性能。

中文摘要 AI 辅助

大型语言模型(LLM)代理越来越多地将自然语言推理与外部工具(如网络搜索和代码执行)交织在一起。这些工具使用策略通常通过强化学习(RL)进行优化,这可能会放大训练数据中的虚假相关性。在这项工作中,我们研究了RL训练的代理何时以及为何学习捷径工具选择策略:即基于表面提示线索而非真实任务需求来调用工具。我们构建了结合事实问答和数学推理任务的受控合成环境,并在训练期间注入与特定工具强相关但对工具必要性因果无关的线索。在反事实评估中,当线索存在但相关工具并非必需时,代理表现出显著的捷径行为,虚假工具调用率增加高达39%。然而,捷径形成并非普遍现象:在我们测试的条件下,只有当代理已经学会可靠地使用目标工具时,捷径才会出现,这表明任务能力而非仅数据集不平衡是捷径学习的关键因素。交换线索分析进一步表明,线索与工具之间的语义对齐显著放大了这一效应。为了缓解这些失败,我们引入了一种密集的、决策级别的奖励,其中LLM评判器评估每次工具调用的必要性。这种工具必要性奖励有效抑制了线索驱动的工具使用,同时保持了任务性能,为提高LLM代理工具使用策略的鲁棒性提供了一种实用方法。

英文摘要

Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the training data. In this work, we study when and why RL-trained agents learn shortcut tool-selection policies: invoking tools based on superficial prompt cues rather than genuine task requirements. We construct controlled synthetic environments combining factual question answering and mathematical reasoning tasks, and inject cues that are strongly correlated with specific tools during training but causally irrelevant to tool necessity. Across counterfactual evaluations where cues are present but the associated tools are not required, agents exhibit substantial shortcut behavior, with spurious tool invocation rates increasing by up to 39 percent. However, shortcut formation is not universal: across the conditions we test, it arises only when the agent has already learned to use the target tool reliably, suggesting that task competence, rather than dataset imbalance alone, is a key factor in shortcut learning. A swapped-cue analysis further shows that semantic alignment between cues and tools substantially amplifies this effect. To mitigate these failures, we introduce a dense, decision-level reward in which an LLM judge evaluates the necessity of each tool call. This tool-necessity reward effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.

发表机构

  • University of Washington(华盛顿大学)
  • University of California San Diego(加州大学圣迭戈分校)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑