模型陈述的拒绝候选人的理由是否起作用?
Does a model's stated reason for rejecting a candidate do any work?
浏览论文内容
中文总结 AI 辅助
本研究通过插入事实句和对照实验,检验语言模型拒绝候选时陈述的理由是否真正影响其选择,发现内容效应存在但受位置和流畅度影响,且测量规则需严格验证。
中文摘要 AI 辅助
当被要求在候选者之间做出选择并解释选择时,语言模型常常通过指出对手档案中缺失的事实来拒绝对手:没有导演,没有死亡日期。该句子是对模型面前文本的断言,无需任何评判者即可进行测试。我们在对手的档案中插入一个真实语料库句子,陈述指定的事实,并在贪婪解码下再次提问。两个对照组将内容与位置分开:在同一档案中插入一个长度匹配的不相关句子,以及在模型从未提及的第三个选项中插入相同的两个句子。在三次运行中最大的一次中,六个开放模型在2WikiMultihopQA上,在模型指名的档案中提供指定事实比不相关对照更能改变其选择,比值比为3.57 [1.54, 8.26],Holm p=0.0210,并且这在剔除任何单个模型后仍然成立。该设计旨在检测的对比,即同一事实出现在无人指名的选项中,未通过校正,Holm p=0.2428。该系列中最强的结果根本不涉及内容断言:相同的不相关句子在指名的对手处比在第三个选项处更能改变选择,Holm p=0.0008。修复和对照在共同候选提及、关系模板和流畅度方面也存在差异;对前两者的事后匹配保留了内容效应的方向,匹配流畅度削弱了一个效应,因此内容对比界定了一个效应而非确立了一个效应。强制单令牌概率读数在相同对比上与自由文本选择的方向不一致,而针对这种不一致的三种候选解释均未找到支持。每项测量都是字符串规则,因此每项都针对其读取的记录进行了验证;验证发现了八个缺陷。最大的一项,即一个选择解析规则在17.1%的可裁决响应中返回了模型刚刚拒绝的选项,本应报告六个存活的对比而非四个。
英文摘要
Asked to choose between candidates and explain the choice, a language model often rejects a rival by naming a fact its profile lacks: no director, no date of death. That sentence is a claim about the text in front of the model, and it can be tested without any judge. We insert a real corpus sentence stating the named fact into the rival's profile and ask again under greedy decoding. Two controls separate content from placement: a length-matched irrelevant sentence at the same profile, and the same two sentences at a third option the model never mentioned. In the largest of three runs, six open models on 2WikiMultihopQA, supplying the named fact at the profile the model named moves its choice more than the irrelevant control does, odds ratio 3.57 [1.54, 8.26], Holm p=0.0210, and this survives dropping any single model. The contrast the design was built to detect, the same fact at the option nobody named, does not clear correction, Holm p=0.2428. The strongest result in the family carries no content claim at all: the identical irrelevant sentence moves the choice more at the named rival than at the third option, Holm p=0.0008. Repair and control also differ in co-candidate mentions, relation template and fluency; post-hoc matching on the first two preserves the content effects' direction, matching fluency weakens one, so the content contrasts bound an effect rather than establish one. A forced single-token probability read disagrees in direction with the free-text choice on that same contrast, and three candidate explanations for the disagreement find no support. Every measurement is a string rule, so each was validated against the records it reads; validation caught eight defects. The largest, a choice-parsing rule that returned the option a model had just rejected in 17.1% of adjudicable responses, would have reported six surviving contrasts instead of four.