arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18919cs.LG

在聚合中迷失:基准测试如何忽视不可替代的模型优势

Lost in Aggregation: How Benchmarks Overlook Irreplaceable Model Strengths

Andrej Tschalzev, Stefan Lüdtke, Heiner Stuckenschmidt, Christian Bartelt

首次发表
浏览论文内容

中文总结 AI 辅助

该研究指出表格机器学习基准的聚合指标忽视模型的数据集特定不可替代优势,提出以数据为中心的峰值性能前沿评估框架,建议基准进展需同时衡量聚合指标与模型对峰值性能集的扩展。

中文摘要 AI 辅助

表格机器学习基准测试通常通过对不同数据集上的分数、排名或 pairwise wins(成对胜率)取平均来总结性能。这类聚合结果对于选择稳健的默认模型很有用,但它们会掩盖一个不同的问题:哪些模型对于在特定数据集上达到峰值性能是必要的?我们认为基准评估还应考虑以数据为中心的峰值性能前沿,该前沿定义为在每个数据集上获得的统计支持最佳的性能。从这个角度来看,模型根据其在前沿上相对于其他模型的位置,可能是不可替代的、充分的、冗余的或易出错的。将此框架应用于 TabArena 基准测试,我们发现常见的聚合指标高度相关,主要衡量一致性和避免失败,而与数据集级别的不可替代性的一致性要低得多。因此,在多个数据集上表现尚可且从未成为最佳选择的模型会受到奖励,而具有独特数据集特定优势的模型在聚合下显得平庸。因此,基准测试的进展不仅应通过聚合指标的改进来衡量,还应通过新模型是否扩展了跨数据集可实现的峰值性能集来衡量。

英文摘要

Tabular machine learning benchmarks typically summarize performance by averaging scores, ranks, or pairwise wins across datasets. Such aggregates are useful for selecting robust default models, but they can obscure a different question: which models are necessary to attain peak performance on particular datasets? We argue that benchmark evaluation should also consider the data-centric peak performance frontier, defined by the best statistically supported performance achieved on each dataset. From this perspective, a model may be irreplaceable, sufficient, redundant, or fallible depending on where it lies on the frontier relative to other models. Applying this framework to the TabArena benchmark, we find that common aggregation metrics are highly correlated and largely measure consistency and avoiding failures, while being much less aligned with dataset-level irreplaceability. Consequently, models performing decently across datasets without ever being the best choice are rewarded while models with unique dataset-specific strengths appear mediocre under aggregation. Hence, benchmark progress should be measured not only by improvements on aggregation metrics but also by whether new models expand the set of attainable peak performances across datasets.

发表机构

  • University of Mannheim(曼海姆大学)
  • University of Rostock(罗斯托克大学)
  • Technical University of Clausthal(克劳斯塔尔工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑