arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33196cs.AI

基准测试可靠吗?通过样本级能力边界进行结构诊断

Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries

Haiquan Hu, Yuzhu Liang, Weicheng Tang, Yanzeng Li, Yao Shi, Tian Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出BSDProbe框架,通过样本级能力边界诊断基准结构,发现基准可靠性因轴而异,并选出区分度提升8.58倍的高价值子集。

中文摘要 AI 辅助

评估大型语言模型(LLM)在很大程度上依赖于基准测试分数,然而聚合指标可能掩盖基准样本是否可靠地支持模型比较这一问题。我们引入了BSDProbe,一个用于基准结构诊断的样本级框架,该框架根据沿有序模型轴的重复杂响应轨迹来估计能力边界。BSDProbe通过边界位置、边界宽度、边界信号有效性和顺序一致性来总结样本,然后将它们聚合成基准级别的结构概况。在六个基准上的实验表明,基准可靠性是轴条件性的且异质的:GSM8K和MATH展现出最稳定的测量结构,MMLU和TriviaQA相对稳定但异质,而GPQA和PopQA则表现出更强的轴条件性风险。这些概况在Qwen3、Qwen2.5和跨模型轴上保持一致。BSDProbe进一步选择紧凑的高价值子集,其模型区分度可达完整基准的8.58倍。这些结果表明,可靠的基准使用需要检查超越排行榜分数的样本级能力边界。

英文摘要

Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boundaries from repeated-response trajectories along ordered model axes. BSDProbe summarizes samples by boundary position, boundary width, boundary-signal validity, and order consistency, then aggregates them into benchmark-level structural profiles. Experiments on six benchmarks show that benchmark reliability is axis-conditioned and heterogeneous: GSM8K and MATH exhibit the most stable measurement structures, MMLU and TriviaQA are relatively stable but heterogeneous, while GPQA and PopQA show stronger axis-conditioned risks. These profiles remain consistent across Qwen3, Qwen2.5, and cross-model axes. BSDProbe further selects compact high-value subsets whose model discriminability reaches up to $8.58\times$ that of the full benchmark. These results suggest that reliable benchmark use requires examining sample-level capability boundaries beyond leaderboard scores.

补充信息

↑