arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

何时可以仅凭评估上线?LLM系统离线评估的试验级替代指标

When Can You Ship on Evals Alone? Trial-Level Surrogacy for Offline Evaluation of LLM Systems

Mårten Schultzberg

arXiv 2610.10142首次发表:更新:

发表机构

Spotify(Spotify(声田))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出试验级替代指标以评估离线评估能否预测在线效应,证明其独立于单位级替代指标,并给出封闭式决策规则及对LLM评判者偏差的分析,实证显示当前评估关系较弱。

AI 中文摘要

离线评估日益指导着基于LLM的产品的变更,但证明评估与人类判断一致,并不能告诉团队哪些变更可以在没有A/B测试的情况下上线。该决策取决于评估处理效应能否预测一类干预措施中的在线处理效应。我们称这一要求为试验级替代指标,并证明它在逻辑上独立于单位级(Prentice)替代指标。在历史发布日志中,两种效应估计中的抽样噪声会衰减观察到的关系。在双变量正态工作模型下,减去报告的干预内方差可恢复干预间协方差,校正后的矩给出了一个封闭形式的上线/测试/放弃规则,该规则限制了上线具有非正向在线效应的变更的概率。LLM评判者增加了另一种扭曲。当评判者在两个臂上同样准确时,评估效应仅被重新缩放,但臂依赖的准确性可能逆转其符号。一旦部署,评估与结果之间的关系在上线和放弃区域中不再可观察,因此如果没有随机审计实验,那里的漂移将变得不可检测。在Upworthy研究档案的183项干预措施中,尽管两侧测量精确,但合并的评估与结果关系较弱,且示例性门控将所有干预措施发送至A/B测试。

英文摘要

Offline evals increasingly guide changes to LLM-powered products, but showing that an eval agrees with human judgment does not tell a team which changes it can ship without an A/B test. That decision depends on whether eval treatment effects predict online treatment effects across a class of interventions. We call this requirement trial-level surrogacy and show that it is logically independent of unit-level (Prentice) surrogacy. In historical launch logs, sampling noise in both effect estimates attenuates the observed relationship. Under a bivariate normal working model, subtracting the reported within-intervention variances recovers the between-intervention covariance, and the corrected moments give a closed-form ship/test/kill rule that bounds the probability of shipping a change with a non-positive online effect. LLM judges add a separate distortion. When the judge is equally accurate on both arms, eval effects are only rescaled, but arm-dependent accuracy can reverse their sign. Once deployed, the eval-to-outcome relationship can no longer be observed in the ship and kill regions, so drift there becomes undetectable without randomized audit experiments. On 183 interventions from the Upworthy Research Archive, the pooled eval-to-outcome relationship is weak despite precise measurement on both sides, and the illustrative gate sends every intervention to an A/B test.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑