DIAL:具有自适应人类偏好校准的位置去偏LLM裁判
DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration
- The University of Hong Kong(香港大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
DIAL提出统一框架,结合LLM与人类比较,分离位置效应并自适应校准偏好,在模拟和基准上验证了稳健性与人类对齐,并贡献了大规模真实数据。
AI中文摘要:
大型语言模型(LLM)作为裁判可实现可扩展的评估,但其判断可能对响应顺序敏感,即使在消除此类位置效应后,仍可能系统性地偏离人类判断。我们引入DIAL,一个统一框架,将丰富的LLM比较与有限的人类比较相结合,以分离裁判特定的位置效应,学习位置去偏LLM偏好中的共享结构,并自适应地将该结构校准至人类偏好目标。理论上,我们研究了DIAL的三个方面:(i)潜在LLM偏好、位置效应和人类校准的识别;(ii)平衡LLM锚定与有限人类证据的自适应估计;以及(iii)校准后人类偏好的固定权重不确定性量化。实证上,我们在受控模拟和三个人类偏好基准上分别评估位置去偏和人类对齐,表明DIAL在响应顺序不平衡时保持稳健,在有限标签下实现强人类对齐排名,并在LLM信息不完美时适应人类证据。我们的真实数据研究收集了来自21个LLM裁判的超过41万条判断,涵盖两种显示顺序,为未来LLM裁判偏差、异质性和人类对齐研究提供了资源。
英文摘要:
Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.