arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分数集之前的核心集:用于大语言模型基准测试的评估无监督提示子集选择

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks

Jihan Yao, Gantavya Bhatt, Arnav Das, Peter Jin, Ke Bao, Qiaolin Yu, Khushi Bhardwaj, Chang Su, Jialei Wang, Yikai Zhu, Sugam Devare, Damon Mosk-Aoyama, Zhen Dong, Venkat Krishna Srinivasan, Yineng Zhang, Oleksii Kuchaiev, Jiantao Jiao, Banghua Zhu, Jeff Bilmes

arXiv 2607.09739首次发表:更新:

发表机构

University of Washington; University of California, Berkeley; Oracle; Together AI; LMSYS; NVIDIA(华盛顿大学; 加利福尼亚大学伯克利分校; 甲骨文公司; Together AI公司; LMSYS; 英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型基准测试的核心集选择,采用评估无监督方法,利用次模子集选择,开发多种次模函数。在新大规模套件上实验发现设施选址函数效果好,该目标不限于特定模式,在相关排行榜上表现优且计算成本低,证明次模性对基准压缩有用。

AI 中文摘要

我们研究大语言模型基准核心集选择问题,即在多个基准测试中选择一小部分提示,使诱导的模型分数和排名接近完整基准测试集的结果。在评估无监督基准核心集选择中,选择算法不使用模型评估结果,通过在多个基准测试中生成提示子集进行细粒度操作。我们使用次模子集选择,并为此开发和评估了许多不同的次模函数。在一个包含35个异构基准测试、18个前沿大语言模型和超61K提示的新大规模套件上,我们发现仅基于廉价语义提示嵌入操作的设施选址函数在一系列核心集预算下比12个基于分数和多样性的基线更好地保留大语言模型分数。此外,我们提出的目标不限于评估无监督模式,在仅需选择少数完整基准测试且有大量模型分数可用的设置中,相同目标在MMLU和MTEB排行榜上与现有最佳基线相当或更优,且计算成本更低。我们的结果表明,一般来说,次模性是基准压缩的强大且可靠工具。

英文摘要

We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and rankings approximate those obtained from the full benchmark suite. In evaluation-unsupervised benchmark coreset selection (our approach), the selection algorithm uses no model evaluation outcomes, and operates on a fine granularity by producing subsets of prompts over multiple benchmarks rather than producing a sub-collection of entire benchmarks. We use submodular subset selection, and we develop and evaluate many different submodular functions for this purpose, including determinantal point process (DPP) based approaches, submodular mutual information functions, and facility location-based functions. On a new large-scale suite of 35 heterogeneous benchmarks spanning five different capability categories, 18 frontier LLMs, and over 61K prompts, we find that the facility location (FL) function operating exclusively on inexpensive semantic prompt embeddings preserves LLM scores better than twelve separate score-based and diversity-based baselines, across a range of coreset budgets. Moreover, we show our proposed objective is not limited to the evaluation-unsupervised regime: in the setting where only a handful of whole benchmarks must be selected and a large amount of model scores are available, the same objective matches or outperforms state-of-the-art baselines on the MMLU and MTEB leaderboards, while being substantially cheaper to compute. Together, our results suggest that submodularity, in general, is a strong and reliable tool for benchmark compression.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑