AI 中文总结
该研究针对以人为中心评估的两难问题,提出AtC两阶段框架,结合人类判断与模型分数,兼具理论保证,在半合成及真实数据集上提升了评估的准确性与鲁棒性。
AI 中文摘要
以人为中心的评估任务是系统决策的关键,高度依赖人类判断,通常缺乏可验证的真实值。现有方法面临两难:仅使用人类判断的方法会受限于异构专业能力和不一致的评分尺度,仅使用模型生成分数的方法则需从不完美的代理或不完整特征中学习。我们提出Aggregate-then-Calibrate(AtC),一个结合两类互补信息源的两阶段框架。第一阶段,使用考虑标注者可靠性的排序聚合模型,将异构比较判断聚合为共识排序;第二阶段,通过对排序的等渗投影校准任何预测模型的分数,在保持模型尽可能多的定量信息的同时,强制序数一致性。理论上,我们证明:(1)建模标注者异质性比同质性能产生严格更高效的共识估计;(2)即使共识排序被误设,等渗校准仍具有风险界;(3)AtC渐近优于仅用模型的评估。在半合成和真实世界数据集上,AtC始终比仅人类或仅模型的评估提升了准确性和鲁棒性。我们的结果将判断聚合与无模型校准相连接,为真实值成本高、稀缺或不可验证时的以人为中心评估提供了原则性方案。
英文摘要
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates any predictive model's scores by an isotonic projection onto the order, enforcing ordinal consistency while preserving as much of the model's quantitative information as possible. Theoretically, we show: (1) modeling annotator heterogeneity yields strictly more efficient consensus estimation than homogeneity; (2) isotonic calibration enjoys risk bounds even when the consensus ranking is misspecified; and (3) AtC asymptotically outperforms model-only assessment. Across semi-synthetic and real-world datasets, AtC consistently improves accuracy and robustness over human-only or model-only assessments. Our results bridge judgment aggregation with model-free calibration, providing a principled recipe for human-centered assessment when ground truth is costly, scarce, or unverifiable.
CommentsAccepted by ICLR 2026