arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14992cs.AIcs.CL

工具结果是否比纯文本更具权威性?针对Claude Opus 5的合成分配任务中虚假主张采纳情况的三项前瞻性研究

Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5

  • Corabo(科拉博公司)

机构由 AI 辅助整理,请以论文原文为准。

Justin Bronder

AI总结:

研究Claude Opus 5在合成任务中,工具结果与纯文本对虚假主张采纳的影响,发现工具结果的权威性未高于已宣布的内联文本。

AI中文摘要:

语言模型系统越来越多地从它们自己也会写入的存储中读取内容,因此,之前仅被写入的主张可能会被当作检索到的内容返回。我们在一项合成查找任务中测试,携带无依据分配任务的消息包是否会改变模型给出的答案。Claude Opus 5会为指定项目选择颜色代码或弃权(不执行)。在一项探索性四臂研究中,当没有目标主张时,虚假代码采纳率为0/24;当之前的助手断言指定了目标时,可评分试验的采纳率为0/22;当工具结果记录指定了目标时,采纳率为14/24;当该结果使用标记为未检查的十字段元数据包装时,采纳率为15/24。工具结果组在12次受支持试验中选择了记录的代码,在24次无依据试验中选择了14次,排除了固定输出 token 偏差,同时留下了大量植入 token 的异质性。一项预先注册的复制研究重现了工具结果与助手断言之间的差距,分别为7/24和0/24,单侧Fisher精确检验p=0.0047。然而,工具结果率在间隔四天进行的试验中从14/24降至7/24。第二项预先注册研究为早期比较设置了实时文本对照:两个记录均提前宣布并置于最终用户回合中,然后在链接的工具结果与后续内联JSON之间交换目标绑定。内联文本在60/60次试验中足以实现虚假代码采纳;工具结果条件下的采纳率为57/60,因此注册的结果优先标准未通过,p=1。该结果并未表明工具结果没有效果,它表明原生工具结果放置并非必要,且本实验未发现结果包比已宣布的内联文本具有更大的行为权重。这些发现涉及单个模型在一个合成任务模板上的表现,通过一个API访问。

英文摘要:

Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.

补充信息

↑