用于评估开放式生成的两级元评分标准:GAMUT,事实完整性基准
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
AI总结:
研究如何评估开放式生成内容的事实完整性,提出两级元评分标准框架并实例化为Gamut基准,构建基于真实图像的问题及评分标准,评估多个模型,发现其具挑战性、区分性且对法官选择鲁棒。
AI中文摘要:
评估长篇生成内容的事实性主要集中在精确性上,即判断模型的陈述是否正确。占主导地位的分解-搜索-验证流程能很好地捕捉错误陈述,但对于回答是否包含应有的所有信息却涉及较少。衡量事实完整性这一事实性缺失的另一半更具挑战性,因为需要枚举完整答案应包含的所有事实,而这些事实很少是简单罗列。我们引入了一个两级元评分标准框架来评估开放式生成,并将其实例化为Gamut,一个长篇生成中事实完整性的基准。该框架基于两级评分标准表示:结构化元评分标准捕捉所需内容的组织和重要性,然后机械编译成一个由二元、机器可评分的评分标准组成的扁平清单,由语言模型法官可靠评分。我们构建了1813个基于10个不同领域真实可穿戴图像的问题,每个问题都配有经专家人工注释验证的有证据支持的评分标准。由于该框架与模态无关,我们还发布了纯文本变体。通过评估14个前沿和开放权重模型,我们发现该基准具有真正的挑战性(Gemini 3.1 Pro的最佳分数为58.7%),具有高度的区分性,并且对法官的选择具有鲁棒性。
英文摘要:
Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric requires expressing the structure of the space of good answers: open-ended sets of acceptable options, ordered processes, and the relative importance of facts. Grading with the rubric requires a judge to score consistently, and judges are far more reliable on flat, binary checks than on rich structure. We resolve this tension with a two-level meta-rubric framework. A structured meta-rubric captures the grading criteria at authoring time, and fixed mechanical rules compile it into a flat checklist of binary, machine-gradable checks that an LLM judge scores reliably at evaluation time. We instantiate the framework as Gamut (Grounded Assessment of Multimodal Factuality), a benchmark for factual completeness in long-form generation. Gamut comprises 1,813 questions grounded in real wearable imagery across 10 diverse domains, each paired with an evidence-backed rubric verified by expert human annotators. Evaluating 14 frontier and open-weight models, we find Gamut genuinely challenging (best score 58.7% from Gemini 3.1 Pro), highly discriminative, and robust to the choice of judge.