arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时停止:LLM评估的贝叶斯最优停止

Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

Toby D. Pilditch

arXiv 2608.14425首次发表:更新:

发表机构

UK AI Security Institute(英国人工智能安全研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出optstop框架,将LLM评估视为序贯测量问题,基于分层贝叶斯推理实现自适应停止,可减少评估试验量且不影响结论,为LLM评估计算资源分配提供新方式。

AI 中文摘要

大语言模型(LLM)评估常采用固定采样预算,对每个项目的采样次数相同,即便估计值已足够精确。我们提出optstop,一种基于精度的自适应停止框架,将评估视为序贯测量问题:在不确定性高的地方继续采样,在估计值足够精确或稳定的地方停止。该框架基于分层贝叶斯推理构建,支持二元、有序和连续结果,且每个基准项目都可参与采样,无需校准项目库。它可实时或回顾性运行,并包含一个保护机制:当测量性能接近零时,采样会更谨慎,而稀有成功最为重要。在一项含200个项目、10个轮次的示例评估中,它在9种验证设置下减少了57%-97%的计划试验,且整体结论与完整运行一致。这些结果表明,LLM评估的计算资源可按不确定性而非固定重复次数分配,节省幅度取决于评估设计。

英文摘要

LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where uncertainty remains high, and stop where estimates are precise or stable enough. The framework builds on hierarchical Bayesian inference, supports binary, ordinal, and continuous outcomes, and keeps every benchmark item eligible for sampling, without requiring a calibrated item bank. It runs live or retrospectively, and includes a safeguard that samples more cautiously as measured performance approaches zero, where rare successes matter most. In an illustrative 200-item, 10-epoch evaluation, it removes 57%-97% of planned trials across nine validation settings, with overall conclusions equivalent to the full run. These results show that LLM evaluation compute can be allocated by uncertainty rather than by fixed repetition counts, with the magnitude of savings depending on evaluation design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑