arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27023cs.LGcs.AIstat.ML

BayesAME:贝叶斯主动模型评估

BayesAME: Bayesian Active Model Evaluation

  • Imperial College London(帝国理工学院)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa

AI总结:

BayesAME是一种序列贝叶斯框架,可自动确定核心集大小以高效评估大型生成模型,其性能优于现有方法,且证实非随机核心集选择及连续响应对数似然的优势。

AI中文摘要:

在不同基准上评估大型生成模型既耗时又计算成本高昂,这催生了对仅通过在项目子集(即核心集)上评估模型来估算完整基准性能的方法的需求。现有文献大多要求从业者输入核心集大小,但当可靠的性能估算优先于效率时,评估方法还应能够自动确定反映该优先级的核心集大小。我们提出BayesAME,一种专门用于自动确定核心集大小的序列贝叶斯框架。BayesAME将性能建模为随机变量,方法是为共享相同历史模型性能的项目组定义潜在能力,联合先验分布编码目标模型与这些历史模型表现相似的信念;这些能力的后验分布用于推导性能估算器、量化性能不确定性,并通过信息增益准则选择添加到核心集的项目,核心集被迭代扩充直至性能估算波动和性能不确定性分别低于用户定义的阈值。我们提出多目标扩展,捕捉多个目标模型间的性能相关性以进一步减小核心集大小。通过在不同基准上的大量实验,我们证明BayesAME始终优于现有方法的序列适配版本;重要的是,我们的综合分析回应了文献中近期的质疑,确立非随机核心集选择比随机选择更具优势;最后,我们强调利用连续响应对数似然而非传统二元评分可显著提升估算准确性。

英文摘要:

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that can estimate full benchmark performance by evaluating models on only a subset of items, known as a coreset. Current literature mostly requires the practitioner to input a coreset size. However, when reliable performance estimation takes priority over efficiency, an evaluation method should also be capable of automatically determining a coreset size that reflects this priority. We introduce BayesAME, a sequential Bayesian framework specifically targeting automatic determination of the coreset size. BayesAME models performance as a random variable by defining a latent ability for each group of items sharing the same historical model performances, with a joint prior distribution encoding the belief that the target model behaves similarly to these historical models. The posterior distribution over these abilities is used to derive performance estimators, quantify performance uncertainty, and select items to add to the coreset via an information-gain criterion. The coreset is iteratively augmented until the performance estimate fluctuation and the performance uncertainty fall below their respective user-defined thresholds. We propose a multi-target extension that captures performance correlations across multiple target models to further reduce the coreset size. Through extensive experiments across diverse benchmarks, we demonstrate that BayesAME consistently outperforms sequential adaptations of existing methods. Crucially, our comprehensive analysis addresses recent skepticism in the literature, establishing that non-random coreset selection is advantageous over random selection. Finally, we highlight that leveraging continuous response log-likelihoods over traditional binary scores significantly enhances estimation accuracy.

↑