AI 中文总结
提出PROOF基准,通过18,486道基于Wikidata的多选题剖析大型语言模型的对象级事实覆盖率,揭示领域差异、方向依赖和干预敏感性,而非单一准确率。
AI 中文摘要
聚合的事实性得分掩盖了语言模型在哪些方面成功、混淆了哪些关系,以及答案是否能经受住对问题或解码器的无关紧要的改变。我们引入了PROOF,一个面向指令微调语言模型事实覆盖率的、基于剖析的基准。PROOF将冻结的Wikidata快照转换为18,486道英文多项选择题,这些题目基于11,779个语义事实、101个类别、392个属性和14个领域。每道题都有一个明确的“我不知道”选项、一个“无正确选项”的控制项,以及九种受控表述;其中1,849道题为无正确选项的陷阱题。我们在166,374个提示上评估了18个开放权重模型部署,并单独对固定的10%子集进行解码扰动。基础事实准确率范围从6.58%到57.59%(随机水平:8.64%),然而每个模型在不同领域间有19.3至36.4个百分点的差距。配对事实揭示了方向相关的检索,通常偏向于主语到宾语查询,但有一个模型的模式发生了反转。在精确分层调整后,我们未发现一致的时间惩罚。中性措辞的改变使准确率变化高达26.5个百分点,而对抗性表述使最初正确的答案最多被破坏79.4%。直接切换到注入的错误标签的变化范围从0.04%到27.5%,表明准确率下降和提示跟随是不同的。选定标记的置信度通常显示出严重的过度自信,解码器扰动使准确率最多移动15.7个百分点,领域剖析移动16.8个百分点。因此,PROOF将事实覆盖率度量为一个结构化的、干预感知的剖析,而不是关于模型“相信”什么的单一论断。
英文摘要
Aggregate factuality scores hide where a language model succeeds, which relations it confuses, and whether an answer survives innocuous changes to the question or decoder. We introduce PROOF, a profile-oriented benchmark for factual coverage in instruction-tuned language models. PROOF converts a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts, 101 classes, 392 properties, and 14 domains. Each question has an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations; 1,849 questions are no-correct-option traps. We evaluate 18 open-weight model deployments on 166,374 prompts each and separately perturb decoding on a fixed 10% subset. Base factual accuracy ranges from 6.58% to 57.59% (chance: 8.64%), yet every model has a 19.3-36.4 percentage-point spread across domains. Paired facts reveal direction-dependent retrieval, usually favoring subject-to-object queries, with the pattern reversing for one model. We find no consistent temporal penalty after exact-stratum adjustment. Neutral wording changes accuracy by as much as 26.5 percentage points, while adversarial formulations break up to 79.4% of answers that were initially correct. Direct switching to an injected false label varies from 0.04% to 27.5%, showing that accuracy loss and hint following are distinct. Selected-token confidence often indicates severe overconfidence, and decoder perturbations move accuracy by up to 15.7 percentage points and domain profiles by 16.8 points. PROOF therefore measures factual coverage as a structured, intervention-aware profile rather than a single claim about what a model "believes."
Comments24 pages, 16 figures, including appendices