发表机构
Johns Hopkins University; University of Illinois Urbana-Champaign; University of California, Los Angeles(约翰·霍普金斯大学; 伊利诺伊大学厄巴纳-香槟分校; 加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨项目反应理论(IRT)用于人工智能评估的可靠性,利用六个大语言模型基准模拟响应矩阵,比较四种估计工具,评估IRT在多方面的表现,发现经典估计器在大型基准中可能不可行,可扩展估计器在特定模型集下可能产生不可靠推断。
AI 中文摘要
人工智能基准越来越多地利用项目级统计模型,特别是项目反应理论(IRT)来估计模型能力、对系统进行排名、选择信息丰富的示例以及诊断基准质量。然而,人工智能基准数据往往与最初开发标准IRT估计工具的人类测试数据模式不同。我们研究了这些模式不匹配如何挑战IRT建模在人工智能评估中的可靠性。通过六个广泛使用的大语言模型基准得出的项目参数和能力分布,模拟三种常见IRT模型下的响应矩阵,并比较四种估计工具。在18000个模拟条件下,系统评估IRT在模型排名、预测性能和项目特征推断方面的计算可行性、可扩展性和可靠性。结果表明,经典估计器在大型基准设置中可能变得不可行,而可扩展估计器在小或非正态分布模型集下可能产生不可靠的项目级和排名推断。本研究确定了潜在特质模型何时可靠支持或可能扭曲人工智能基准测试声明,以及需要何种样本量和诊断才能可靠使用。
英文摘要
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or non-normally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.