arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预测驱动的平滑与验证用于细粒度AI评估

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Sho Kawano, Zehang Richard Li, Paul A. Parker

arXiv 2609.20758首次发表:更新:

发表机构

University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI细粒度评估中标签稀少导致直接估计不精确的问题,提出预测驱动平滑(PP-S)及其分类体系扩展(PP-TS)的贝叶斯估计方法,并推导近似无偏的交叉验证得分,在基准和真实智能体流量上显著提升点/区间估计精度与选择可靠性。

AI 中文摘要

评估AI系统需要细粒度(disaggregated)的评估,因为性能在不同领域(如基准测试任务类型或部署智能体中的对话类型)之间存在差异。穷举测试成本高昂,因此评估依赖于对已标注单元样本的抽样。我们将评估集视为有限总体,并寻求对每个领域均值的精确点估计和区间估计。直接估计方法(包括预测驱动推断(PPI))仅使用领域自身的标签,在标签稀少时精度不足。小域估计(small area estimation)解决了这一问题,我们在此基础上构建了一个集成的估计与验证工作流程。在估计方面,我们提出了预测驱动平滑(PP-S),这是一种对每个领域的预测驱动估计进行拟合的贝叶斯模型,并扩展了跨报告分类体系(taxonomy)借用信息强度的版本(PP-TS)。在验证方面,我们推导了一种新的、近似无偏的基于设计的交叉验证得分,用于在直接估计器和平滑估计器之间进行选择。我们研究了一个具有可验证评分的精选基准,以及由人类评分的部署智能体流量,两者均观测到了所有结果。在两种情况下,所提出的估计器在点估计和区间估计上均优于直接估计器,且覆盖率接近名义水平。在相同的抽样预算下,我们的得分选择效果与独立验证样本相当,并且对所选估计器误差的估计远更准确。

英文摘要

Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, so evaluation rests on a sample of labeled units. We treat the evaluation set as a finite population and seek accurate point and interval estimates of each domain mean. Direct estimators, including prediction-powered inference (PPI), use only a domain's own labels and are imprecise where labels are few. Small area estimation addresses this problem, and we build on it to develop an integrated workflow for estimation and validation. For estimation, we propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, with an extension that borrows strength across a reporting taxonomy (PP-TS). For validation, we derive a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators. We study a curated benchmark with verifiable grading and deployed agent traffic graded by humans, each with every outcome observed. In both, the proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage. At the same sampling budget, our score selects as well as an independent validation sample does and estimates the selected estimator's error far more accurately.

Comments15 pages of main text, 30 pages total, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑