AI基准测试实际测量什么:将收敛效度与判别效度适配于对五十六个AI基准的审视
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
- University of Michigan(密歇根大学)
- Stanford University(斯坦福大学)
- Microsoft Research(微软研究院)
- Abridge
- Yale University(耶鲁大学)
- Cornell Tech(康奈尔科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究将收敛效度与判别效度方法应用于56个AI基准,发现安全与能力概念在不同基准间测量不一致,部分基准实际测量内容与声称不符。
AI中文摘要:
基准测试在模型开发与治理中扮演核心角色,然而,这些基准是否真正测量了其声称要测量的概念(如推理能力、拒答能力)往往并不明确。我们将社会科学中的收敛效度与判别效度方法适配为一种审视AI基准的途径,并将其应用于涵盖53个模型的56个能力与安全基准。我们将声称概念实质相似的基准归入同一指定概念,并考察同一指定概念下基准间的模型排名相关性是否强于不同指定概念下基准间的排名相关性。我们利用项目反应理论(IRT)模型在项目层面提出类似问题。我们发现,同一指定安全概念下基准间的模型排名相关性往往较弱,这表明这些概念在不同基准间的概念化可能不一致。对于指定的能力概念(如推理、知识),同一指定概念下基准间的模型排名相关性往往与不同指定概念下基准间的相关性一样强,这表明这些能力概念可能彼此区分度不高。在某些情况下,共享设计元素(如评分格式)的基准之间的相关性比同一指定概念下的基准更强。最后,某些个别基准与指定为不同概念的基准的相关性,强于与其自身指定概念相同的基准,这表明它们可能测量的并非其声称的概念。例如,BBQ-accuracy与标记为推理的基准的相关性,强于与其自身指定概念(偏见)相同的基准。为支持未来关于基准效度的实证研究,我们发布了在项目层面和基准层面的广泛模型输出与得分数据集。
英文摘要:
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.