arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机大语言模型的投机式评估

Speculative Evaluation of Stochastic LLMs

Qianli Shen, Xiang Li, Ruomeng Ding, Yanxi Chen, Daoyuan Chen, Yaliang Li

arXiv 2609.28560首次发表:更新:

发表机构

Alibaba Group; National University of Singapore; University of North Carolina at Chapel Hill(阿里巴巴集团; 新加坡国立大学; 北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对随机大语言模型评估成本高的问题,提出基于分层贝叶斯奈曼分配的投机式评估方法,通过试点和异步执行降低方差,在多个基准上平均减少12.8%-33.6%的方差。

AI 中文摘要

评估随机大语言模型成本高昂:基准分数通过随机回放估计预期性能,但均匀重复忽略了任务级回放方差的显著差异。我们研究如何在精确的回放预算下最小化固定基准均值的方差。我们开发了投机式评估方法,采用分层贝叶斯奈曼(HBN)策略,其中试点规模和阶段权重事先联合选择。该方法运行一个短均匀试点,用分层贝叶斯模型汇总每个任务的成功计数,并利用任务级采样方差的后验期望进行精确正整数奈曼分配。为缓解试点同步障碍,HBN-async 从部分试点反馈中投机式执行延续,并保留最终分配所选的那些。在六个检查点和18个基准组上,我们评估了107个非退化基准-检查点配置。对于每个任务8-64的回放预算,投机式评估相对于均匀方法在配置上平均降低方差12.8%-33.6%,优于事后调整的经验和独立贝叶斯基线。考虑试点同步障碍的真实生成实验表明,HBN-async 缓解了其开销,有助于将统计效率转化为实际评估效益。

英文摘要

Evaluating a stochastic large language model is costly: benchmark scores estimate expected performance from randomized rollouts, yet uniform repetition ignores sharp differences in task-level rollout variance. We ask how to minimize the variance of a fixed-benchmark mean under an exact rollout budget. We develop Speculative Evaluation with a Hierarchical Bayesian Neyman (HBN) policy with pilot size and stage weight jointly chosen ex ante. It runs a short uniform pilot, pools per-task success counts with a hierarchical Bayesian model, and uses posterior expectations of task-level sampling variances for exact positive-integer Neyman allocation. To mitigate the pilot synchronization barrier, HBN-async speculatively executes continuations from partial pilot feedback and retains those selected by the final allocation. Across six checkpoints and 18 benchmark groups, we evaluate 107 nondegenerate benchmark-checkpoint profiles. For rollout budgets of 8-64 per task, Speculative Evaluation reduces variance relative to Uniform by 12.8%-33.6% on average across profiles, outperforming hindsight-tuned empirical and independent Bayesian baselines. Real-generation experiments that account for the pilot synchronization barrier show that HBN-async mitigates its overhead, helping translate statistical efficiency into practical evaluation benefits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑