arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25645cs.LGcs.CLstat.ML

通过贝叶斯Bandit Gittins指数实现高效成本感知的LLM评估

Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices

Qian Xie, Yueli He, Nairen Cao

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM配置评估成本高的问题,提出基于贝叶斯Bandit Gittins指数的GittinsEval方法,通过最优停止和LCB评分实现高效评估,仅用1%-2%成本即达近零遗憾。

中文摘要 AI 辅助

为了识别高性能的候选大语言模型(LLM)配置,对每个基准测试项上的每个候选配置进行详尽评估成本高昂。我们将配置选择问题建模为成本感知的贝叶斯Bandit问题,并提出GittinsEval,该方法利用贝叶斯最优的Gittins策略来决定接下来评估哪个配置以及何时停止。我们通过一种随时推荐规则扩展了该策略,该规则同时适用于完全评估和部分评估的配置,使用LCB风格的分数来考虑后验不确定性。GittinsEval计算效率高,在离线预计算后仅需轻量级的在线更新。在GSM8K、PIQA、AlpacaEval和MMLU响应矩阵上,GittinsEval始终具有竞争力,尤其是在大型示例基准上相对于配置级贝叶斯优化,以及在大型候选任务上相对于不考虑成本的Bandit基线,取得了显著优势。关键的是,GittinsEval通常仅使用详尽评估成本的1%至2%即可实现接近零的简单遗憾;它还提供了一种自适应停止规则,通常在1%至10%时触发。

英文摘要

Exhaustively evaluating every candidate LLM configuration on every benchmark item to identify a high-performing one is costly. We formulate configuration selection as a cost-aware Bayesian bandit problem and propose GittinsEval, which draws on the Bayesian-optimal Gittins policy to determine which configuration to evaluate next and when to stop. We extend the policy with an anytime recommendation rule over both fully and partially evaluated configurations, using an LCB-style score to account for posterior uncertainty. GittinsEval is computationally efficient, requiring only lightweight online updates after offline precomputation. Across GSM8K, PIQA, AlpacaEval, and MMLU response matrices, GittinsEval is consistently competitive, with particularly strong gains over configuration-level Bayesian optimization on large-example benchmarks and over cost-unaware bandit baselines on large-candidate tasks. Crucially, GittinsEval often attains near-zero simple regret using only 1% to 2% of the exhaustive-evaluation cost; it also offers an adaptive stopping rule that typically triggers at 1% to 10%.

发表机构

  • Cornell University(康奈尔大学)
  • Columbia University(哥伦比亚大学)
  • New York University(纽约大学)
  • Shanghai University of Finance and Economics(上海财经大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑