arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02966cs.CL

每个错误答案都有价值:面向LLM多项选择基准的选项层面心理测量学

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

Xiao Fei, Yang Zhang, Sarah Almeida Carneiro, Michalis Vazirgiannis

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM多项选择基准提出LLM-NRM框架,利用错误答案的独特信息提升能力估计精度与基准测试效率,其能力估计与人类偏好排行榜相关性达0.920,仅错误回答就能实现0.943的高相关。

中文摘要 AI 辅助

大多数多项选择题(MCQ)基准仅通过大型语言模型(LLM)是否选择正确答案来评估它们。这种二元评分将所有错误回答同等对待,即便LLM在错误选项中的偏好可能包含关于其行为和能力的系统且有用的信息。我们引入LLM名义响应模型(LLM-NRM),这是一个感知选项的心理测量框架,它对所有答案选项的完整分布进行建模,以联合估计LLM能力和选项层面的题目特征,同时分离出特定于模型的响应校准锐度、位置偏好以及依赖难度的弃权(不执行)行为。在来自14个基准的189个LLM和31554个题目上,LLM-NRM对留存的LLM-题目交互的预测比二元题目响应模型和传统名义响应基线更准确,其能力估计与外部人类偏好Elo排行榜的斯皮尔曼相关系数达到0.920,表现最强。干扰项身份相比正确性,每个题目额外贡献+101%的费希尔信息,仅错误回答就能恢复完整信息的能力估计,斯皮尔曼相关系数为0.943。学习到的题目参数还支持高效基准测试,其中41个选定题目以肯德尔相关系数0.85保留了全题库排名,对应减少了770倍。综上,我们证明错误答案承载独特且有用的测量信息,而非代表等价的错误。

英文摘要

Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.

↑