arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BACON:用于多人工智能评判器建模与评估的预算人类校准

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha

arXiv 2607.16239首次发表:更新:

发表机构

Adobe Research; University of California, Berkeley(Adobe研究院; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对人工智能评判器输出有偏差等问题,提出BACON四阶段流程,结合预算人类校准与多人工智能评判器输出。通过构建辅助特征、收集人类标签训练模型,实现总体指标估计和个体级替代评分,提高了预测准确性等,提供实用评估框架。

AI 中文摘要

人工智能评判器为人工评估提供了一种可扩展、低成本的替代方案,但其输出可能存在相对于人类偏好的偏差,且高度依赖项目,在评判器、任务和领域之间存在差异。当未校准的人工智能评估用于模型排名、项目评分或总体质量报告时,这些偏差会直接扭曲下游决策。我们提出了BACON,这是一个四阶段的流程,将预算人类校准与多个人工智能评判器的输出相结合,以产生更准确的注释。BACON为每个项目构建全覆盖的辅助特征,包括多评判器分数、令牌级不确定性统计和上下文嵌入。然后,它为一个小的采样子集收集人类标签,并训练一个交叉拟合的结果模型,以生成校准后的项目级替代预测。这些预测支持两个用例:使用具有有效置信区间的增强估计方程估计器对总体指标(如均值或分位数)进行总体估计;以及用于项目排名和注释的个体级替代评分。BACON将人工智能评判器视为辅助测量而非地面真值:人类标签提供校准锚点,而人工智能衍生的信号提高效率。在不同的任务、领域和标注预算中,BACON提高了预测准确性和排名一致性,并相对于原始人工智能输出和基于纯人类标签的方法减少了偏差和方差。这些结果表明,BACON为有限人工注释的可扩展评估提供了一个实用的、基于统计的框架。

英文摘要

AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or population-level quality reporting, these biases can directly distort downstream decisions. We propose BACON, a four-stage pipeline that combines budgeted human calibration with multiple AI-judge outputs to produce more accurate annotations. BACON constructs full-coverage auxiliary features for every item, including multi-judge scores, token-level uncertainty statistics, and contextual embeddings. It then collects human labels for a small sampled subset and trains a cross-fitted outcome model to generate calibrated item-level surrogate predictions. These predictions support two use cases: population-level estimation of summary metrics, such as means or quantiles, using an augmented estimating-equation estimator with valid confidence intervals; and individual-level surrogate scoring for item ranking and annotation. BACON treats AI judges as auxiliary measurements rather than ground truth: human labels provide the calibration anchor, while AI-derived signals improve efficiency. Across diverse tasks, domains, and labeling budgets, BACON improves predictive accuracy and ranking consistency, and reduces bias and variance relative to raw AI outputs and purely human-label-based methods. These results show that BACON offers a practical, statistically grounded framework for scalable evaluation with limited human annotation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑