AI 中文总结
该研究将事实回忆分解为生成、排序和选择三步,发现显式验证($P(\mathrm{True})$)相比生成式似然能显著提升候选排序准确率(AUROC 提升 0.08–0.12)和多数票准确率(约 5 点),且收益受实体可及性和正确性定义影响。
AI 中文摘要
生成正确答案并不意味着语言模型会选择它。我们将事实回忆分为三个步骤:生成正确候选、对可用候选进行排序以及选择最终答案。生成前读取(pre-generation readouts)能预测事实回忆以及三个模型家族中采样将覆盖哪些问题,但对于可用的正确答案最终是否会被选中却鲜有说明。使用 $P(\mathrm{True})$ 进行显式验证在 Gemma、Qwen3 和 Llama 中改善了问题内排序,相比平均对数似然,AUROC 提升了 $0.08$--$0.12$。在一个前瞻性定义的 Gemma 队列中,验证将多数票准确率提高了约 5 个百分点,并且相比聊天模板似然(一个更强的生成基线)仍能获得约 2 个百分点的提升。该优势在具有常见答案先验的关系上最为显著,并且依赖于对实体的访问;遮蔽实体在较大的 Qwen 模型中消除了排序优势。最后,测量到的收益取决于正确性的定义方式:面向回忆的参考匹配可能将选项列表归功于似然,从而大幅低估在人类语义判断下所见的改进。
英文摘要
Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with $P(\mathrm{True})$ improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of $0.08$--$0.12$. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about $5$ points, and still gains about $2$ points over chat-template likelihood, a stronger generative baseline. The advantage is strongest for relations with common-answer priors and depends on access to the entity; masking the entity removes the ranking advantage in larger Qwen models. Finally, the measured benefit depends on how correctness is defined: recall-oriented reference matching can credit option lists favored by likelihood and substantially understate the improvement seen under human semantic judgments. Prior work shows that models can carry latent factual knowledge and judge candidate answers; we show that these capabilities do not collapse into a single notion of ``knowing,'' and trace where information is gained, lost, or mismeasured between availability, ranking, and final choice.