发表机构
UWCSEA East Campus; Haileybury Astana; The Shishukunj International School; Apta AI; Spark AI Research(东南亚联合世界学院东校区; 黑利伯瑞阿斯塔纳学校; 希舒昆杰国际学校; Apta AI; Spark AI Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究不必要的工具可用性如何降低LLM基于自身知识回答的准确率,并发现一句范围感知指令可恢复大部分性能损失。
AI 中文摘要
大型语言模型(LLMs)越来越多地部署了外部工具,这些工具扩展了它们超越自身知识的能力。工具在需要外部信息的任务上有所帮助,但它们的可用性也可能改变模型处理不需要这些工具的问题的方式。先前的工作主要询问模型是否适当地选择和使用了工具;而不必要的工具是否会改变答案的正确性则较少受到关注。我们探究了提供一个相关但不必要的工具是否会影响模型从自身知识回答的能力,以及先前的工具交互是否会改变这一行为。我们在10个知识领域构建了500对查询。每对包含一个工具查询(需要该领域的工具)和一个封闭域查询(不需要工具)。我们评估了六个LLM,分别在工具不可用、可用以及在先前工具调用后可用的情况下进行测试。在3,000次基线试验中,汇总的回答率为98.2%。当不必要的工具可用时,该比率降至63.5%,且模型之间存在较大差异。即使工具很少被调用,这种下降也会发生,因此不能仅用不必要的工具调用来解释。一句范围感知的系统指令可以恢复大部分丢失的答案。
英文摘要
Large language models (LLMs) are increasingly deployed with external tools that extend what they can do beyond their own knowledge. Tools help on tasks that need external information, but their availability may also change how a model handles questions that do not need them. Prior work has mostly asked whether models select and use tools appropriately; whether an unnecessary tool changes the correctness of answers has received less attention. We ask whether making a related but unnecessary tool available affects a model's ability to answer from its own knowledge, and whether a preceding tool interaction changes this behaviour. We construct 500 query pairs across 10 knowledge domains. Each pair consists of a tool query, which needs the domain's tool, and a closed-domain query, which does not. Six LLMs are evaluated with the tool unavailable, available, and available after a prior tool call. Across 3,000 baseline trials the pooled answer rate is 98.2%. When an unnecessary tool is available it falls to 63.5%, with large differences between models. The decrease occurs even when the tool is rarely called, so it cannot be explained by unnecessary tool invocation alone. A one-sentence scope-aware system instruction recovers most of the lost answers.
Comments10 pages, 2 figures, 4 tables