发表机构
Corabo(科拉博公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究分析了配备工具的语言模型作出未获支持最终主张的问题,在 Qwen3-32B 上实现了 33/512 的发生率,提供证据可修复全部此类主张,而 Gemma 4 未出现该问题。
AI 中文摘要
配备工具的语言模型可能会在未获得所见证据支持的情况下作出最终主张,即便存在单个可用工具调用即可解决不确定性,且其指令明确禁止假设与猜测。我们将此失败拆分为两个精确定义的量:发生率(模型自行作出未获支持主张的频率,基于可见证据与最终主张测量,不使用隐藏的正确答案);以及条件修复率(当提供缺失证据时,这些自然产生的未获支持主张被修复的频率)。在一个固定的 Qwen3-32B 设置中,对 256 个新提示模板的 512 次首次响应里,有 33 次以未获支持的已确立主张结束。我们从主张发生时的精确状态副本中重放每个案例;在每个匹配的重放中,替代工具响应具有相同的结构和长度,仅在一个字符的响应代码上存在差异。提供证据修复了全部 33 次主张;携带无用信息的匹配响应未修复任何一次。当证据支持原始答案时,模型保留全部 33 次,未观察到任何损害。在另一项实验中,在需要证据的 64 个案例里,自动检查规则添加了 21 次证据调用,纠正了全部 10 次错误的未获支持主张,保留了 11 次偶然正确的主张,且从未将正确答案改为错误答案。在使用相同采样设置的固定 Gemma 4 设置中,模型在全部 512 次首次响应中都调用了工具,从未作出未获支持的最终主张,因此无法测量该设置的条件修复率。这些结果描述了两个合成任务系列中两个局部固定模型设置的情况,未表明此失败在实际部署中的常见程度,也未表明其反映了模型间共享的通用机制。
英文摘要
A language model with access to tools can commit to a final claim unsupported by the evidence it has seen, even when a single available tool call would resolve the uncertainty and its instructions explicitly forbid assumptions and guesses. We separate this failure into two precisely defined quantities: occurrence, how often the model makes an unsupported claim on its own, measured from the visible evidence and final claim without using the hidden correct answer; and conditional repair, how often those same naturally occurring unsupported claims are repaired when the missing evidence is supplied. On one fixed Qwen3-32B setup, 33 of 512 first responses to 256 new prompt templates ended with an unsupported established claim. We replayed each case from an exact copy of the state in which the claim occurred; within each matched replay, the alternative tool responses had the same structure and length and differed only in a one-character response code. Resolving evidence repaired 33 of 33 claims; a matched response carrying no useful information repaired 0 of 33. When the evidence supported the original answer, the model preserved 33 of 33, with no observed harm. In a separate experiment, on 64 cases where evidence was needed, an automatic checking rule added 21 evidence calls, corrected all 10 wrong unsupported claims, preserved the 11 that were correct by accident, and never changed a correct answer into a wrong one. On a fixed Gemma 4 setup using the same sampling settings, the model called the tool in all 512 first responses and never made an unsupported final claim, so conditional repair could not be measured for that setup. These results describe two local fixed model setups on two synthetic task families. They do not show how common this failure is in real-world deployments, nor that it reflects a general mechanism shared across models.
Comments18 page, 1 figure, 5 tables