arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03340cs.CL

基准测试基准:测试常识基准的预测有效性

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

Ine Gevers, Walter Daelemans

AI总结:

该研究评估了23种模型在多类常识基准、对照任务及下游推理任务上的表现,发现常识基准仅对少数下游任务有一致预测性,修改基准未提升预测能力,其仅提供任务相关的下游常识能力证据。

AI中文摘要:

预测大语言模型(LLM)在现实世界任务上的能力至关重要,但常识基准上的性能能在多大程度上预测下游性能仍未明确。为确定广泛采用的常识基准的实际可用性,我们评估了6个系列的23种模型,涉及4个既定常识基准、4个修改变体、3个非常识对照任务,以及8个需要隐性社会、语用、时间或物理推理的下游任务。我们比较模型排名、计算受控相关性,并使用留一模型系列交叉验证来评估常识基准的效标效度。结果显示,修改后的基准在很大程度上保留了原始模型排名,且未提升下游预测能力。常识基准仅对少数下游任务表现出跨系列一致的预测有效性,其余任务收益较小或与特定指标相关。总体而言,标准化常识基准提供的是与任务相关的证据,而非对下游常识能力的广泛证明。

英文摘要:

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.

↑