发表机构
Cimat(西马研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型基准神经项目反应理论模型,提出拉普拉斯-PSN-IRT方法,通过近似贝叶斯后验推断增强模型,能恢复校准不确定性,可进行可信区间等分析,实验表明该方法在模型比较、信息稳定性及能力排名恢复等方面表现良好。
AI 中文摘要
项目反应理论(IRT)最近被提出作为评估大语言模型(LLM)基准的框架,可将模型潜在能力与单个基准项目属性分离。现有神经IRT方法,包括PSN-IRT,使用点估计来估计这些量,限制了不确定性量化和下游统计推断。我们引入拉普拉斯-PSN-IRT,一种事后最后一层拉普拉斯近似,通过近似贝叶斯后验推断增强训练好的PSN-IRT模型,无需重新训练即可恢复模型能力和项目难度的校准不确定性。由此产生的后验使得可信区间、模型间概率比较以及参数不确定性传播到基于费舍尔信息的项目选择成为可能。我们表明,在标准LLM基准排行榜上的12个模型之间,尽管点估计排名不同,但大多数成对比较在统计上没有显著差异。我们还表明,由于点估计费舍尔信息是在单个参考能力下评估的,许多基准项目的点估计费舍尔信息可能几乎为零,而后验期望费舍尔信息在整个能力范围内保持更稳定。最后,在大多数实验设置中,后验期望费舍尔信息能更准确地从小基准子集中恢复全基准能力排名,同时在最小子集上与点估计性能相匹配。我们使用留出的预测覆盖率验证了近似后验的校准,并发现将项目难度建模为随机而将项目区分度建模为固定在该架构中产生了校准良好的不确定性。
英文摘要
Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty quantification and downstream statistical inference. We introduce Laplace-PSN-IRT, a post-hoc last-layer Laplace approximation that augments a trained PSN-IRT model with approximate Bayesian posterior inference, recovering calibrated uncertainty over model ability and item difficulty without retraining. The resulting posterior enables credible intervals, probabilistic comparisons between models, and propagation of parameter uncertainty into Fisher-information-based item selection. We show that most pairwise comparisons among 12 models on a standard LLM benchmark leaderboard are not statistically distinguishable despite differing point-estimate ranks. We further show that point-estimate Fisher information can become nearly zero for many benchmark items because it is evaluated at a single reference ability, whereas posterior-expected Fisher information remains substantially more stable across the ability range. Finally, posterior-expected Fisher information more accurately recovers full-benchmark ability rankings from small benchmark subsets in most experimental settings while matching point-estimate performance for the smallest subsets. We validate the calibration of the approximate posterior using held-out predictive coverage and find that modeling item difficulty as random while treating item discrimination as fixed produces well-calibrated uncertainty in this architecture.