发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估29个LLM在HLE多项选择子集上的表现,发现HLE仅测量单一一般推理因素,其领域子分数不具区分能力,且对前沿模型的区分能力有限。
AI 中文摘要
人类终极考试(Humanity's Last Exam, HLE)被广泛用于评估前沿语言模型,该考试将问题分为8个学科领域类别,其子分数常被解读为不同能力的证据。然而,尚无研究评估这些标签是否对应经验上可分离的潜在构念,也未评估该基准是否能有效区分能力相近的模型。我们在HLE仅文本的多项选择子集(J=428个题目)上评估了29个大语言模型(LLM),并运用心理测量方法评估该基准的维度及其测量精度分布。拟合两参数逻辑斯蒂项目反应理论(IRT)模型后,我们发现一致证据表明HLE测量单一的一般推理因素:McDonald's ω_h=0.998,领域标签仅能解释3.5%的项目反应方差,领域内与领域间的残差相关几乎相同(Cohen's d=0.016),且领域特定能力估计与总分近乎冗余(r≥0.81)。对测验信息函数的单独分析显示,测量精度集中在中等能力水平,在前沿模型所处的θ=0以上则急剧下降。这些发现表明,HLE的领域子分数不能作为不同能力解读的依据,且该基准区分最强模型的能力有限。
英文摘要
Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find that HLE is dominated by a single general reasoning factor: domain labels explain only 3.5\% of variance in the leading principal components of item responses, and within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$). A simulation study with the same number of models ($N = 29$) shows that our analyses would detect clearly distinct domain abilities, so their absence is informative; however, small domain-specific differences cannot be ruled out, and model rankings vary across domains somewhat more than a single factor predicts. A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $θ= 0$, where the strongest models sit. These findings suggest that the subset's domain subscores do not warrant distinct capability interpretations and that its ability to discriminate among the strongest models might be limited.