AI 中文总结
本研究揭示受限选项评分在模型即将输出非选项内容时产生误导性分数,提出无标签诊断方法并验证预填充答案主干可显著提升预测性能。
AI 中文摘要
受限选项评分读取模型对一组固定允许答案的概率,因此即使模型即将写出其他内容,它也会返回一个分数。我们研究了一个提示,该提示引用了一个多项选择题,以题目自身的答案指令结尾,并要求预测仅给定简短说明的读者模型是否能正确回答该问题。模型可以开始回答所引用的题目:在三个Qwen3检查点上,答案字母在几乎每个段落中都是最可能的词元,并且预测将读者的正确性排名与随机水平无法区分。除了先前工作中的argmax检查和预填充修复之外,我们贡献了一种无标签的诊断方法,用于检测声明选项的支持度,即它们在上下文中的首词元质量;一个删除测试,将Qwen3答案字母的开头追溯到所引用的指令;以及渲染和分词检查,用于检测丢失的分隔符和重合的首词元。预填充答案主干将中位数选项质量从低于提高到高于一半,在所有使用此提示评分的九个检查点上,并且在选项和重新归一化不变的情况下,将Qwen3平均AUC从0.4929提高到0.6152。lm-polygraph的默认P(True)估计器显示出相关的失败:它在Qwen3不思考时偏向答案字母(MMLU)或提示的(A)/(B)标签(TriviaQA)的位置读取True,并且其在MMLU上的Qwen3平均AUROC与随机水平无法区分,在TriviaQA上低于随机水平。在答案主干之后对(A)/(B)标签进行评分和重新归一化将其提高到0.632和0.868,就地针对False重新归一化True达到0.660和0.860,并且在TriviaQA上两者都超过了所有十四个默认单答案库估计器。原始预测也被重新归一化,但仍然与随机水平无法区分:选项质量在两种设置中无需标签即可标记低支持度,只有检查预期目标才能显示哪个分数仍然排名。
英文摘要
Constrained-option scoring reads a model's probabilities for a fixed set of permitted answers, so it returns a score even when the model is about to write something else. We study a prompt that quotes a multiple-choice item, ending in the item's own answer instruction, and asks for a forecast of whether a reader model given only a short note will answer it correctly. The model can begin answering the quoted item: on three Qwen3 checkpoints an answer letter is the most probable token on nearly every passage, and the forecast ranks the reader's correctness indistinguishably from chance. Beyond prior work's argmax check and prefill repair, we contribute a label-free diagnostic of the declared options' support, their first-token mass in context; a deletion test tracing the Qwen3 answer-letter start to the quoted instruction; and rendering and tokenisation checks for dropped separators and coinciding first tokens. Prefilling an answer stem lifts the median option mass from below to above one half on all nine checkpoints scored with this prompt and, with options and renormalisation unchanged, raises the Qwen3-averaged AUC from 0.4929 to 0.6152. lm-polygraph's default P(True) estimator shows a related failure: it reads True where Qwen3 without thinking favours an answer letter (MMLU) or the prompt's (A)/(B) label (TriviaQA), and its Qwen3-averaged AUROC is indistinguishable from chance on MMLU and below chance on TriviaQA. Scoring and renormalising the (A)/(B) labels after an answer stem raises it to 0.632 and 0.868, renormalising True against False in place reaches 0.660 and 0.860, and on TriviaQA both exceed all fourteen default single-answer library estimators. The original forecast is renormalised too, yet stays indistinguishable from chance: option mass flags low support in both settings without labels, and only checking the intended target shows which score still ranks.