arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准平衡:面向基准多样性与任务条件评估的语义密度重加权

Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison

Jhen-Ke Lin, Hong-Yun Lin

arXiv 2608.30044首次发表:更新:

发表机构

National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出BoB方法,通过逆密度语义权重与残差场,实现对基准多样性的鲁棒评估与任务条件排名,提升了模型评估的准确性与可控性。

AI 中文摘要

语言模型通常通过对基准列表的分数取平均值来进行比较,且所有基准权重相等。这类基准列表是在未明确测量设计的情况下通过发表内容逐步扩充的,因此等权重会将已发表基准的密度转化为隐含的能力权重:被密集基准测试的区域会被重复计数。我们提出了基准平衡(Balance of Benchmarks, BoB)方法,该方法嵌入基准描述并为每个基准分配逆密度语义权重,邻近条目会在公开的密度尺度上共享聚合影响。在将异构分数等效到共同潜在尺度后,残差场会使用相同的几何结构根据任务查询对模型排名进行条件调整,这两个组件具有不同的经验作用。在包含586个模型和14个基准的快照中,BoB能够预测哪些模型在保留的任务上表现异常出色(超出其一般能力),其轮廓相关系数达到0.462,而等权重下仅为0.049;它还限制了密集重复基准对聚合结果的影响,在依次添加每个基准的4个副本后,生成的排名保留了Kendall tau系数0.995,而等权重下仅为0.936。因此,残差场提供了任务条件预测能力,逆密度权重则提升了对基准多样性的鲁棒性,二者共同将基准列表的构成从评估套件的偶然属性转化为测量设计中明确可控的部分,为面向任务且对基准多样性鲁棒的模型评估提供了原则性基础。

英文摘要

Model comparison increasingly relies on large collections of publicly reported benchmark scores, yet common aggregation strategies trade off evidence coverage against control over capability weighting. Manually curated suites leave potentially informative evaluations unused, while uniform averaging retains them but gives greater influence to capabilities that happen to be benchmarked more densely. We introduce Balance of Benchmarks (BoB), a framework that retains eligible benchmark evidence while adapting its influence for task-conditioned model comparison using only public aggregate scores. BoB combines semantic density weighting, score equating across benchmarks of different difficulty, and task-relevant residual pooling. We evaluate it on 605 configurations across 14 Artificial Analysis benchmarks and on WildScores, a collection of 148 developer-reported benchmarks evaluated with held-out source-lineage families. On WildScores, BoB-Support raises family-mean Spearman correlation from 0.764 under uniform standardized averaging to 0.823, reduces MAE from 6.19 to 5.10 normalized score points, and increases three-model shortlist hit rate from 65.3% to 72.6%. BoB-Constant reaches a Spearman correlation of 0.831 and a hit rate of 74.6%. Separately, density weighting reduces average ranking changes when benchmarks are repeated, including as paraphrased copies. BoB-Support also reduces retrospective three-model shortlist regret from 2.08 to 1.67 normalized score points. BoB makes benchmark inclusion, redundancy, and task relevance explicit and testable measurement choices, allowing existing benchmark evidence to be used more fully while moderating the influence of benchmark proliferation.

Comments65 pages including references and appendices. Expanded evaluation with WildScores, a collection of 148 developer-reported benchmarks

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑