发表机构
Beihang University; Zhongguancun Academy; Hangzhou Innovation Institute of Beihang University(北京航空航天大学; 中关村学院; 北京航空航天大学杭州创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对有偏LLM评判者的主动偏好学习,提出扰动调整最优设计(NAOD),通过优先选择策略相关信息并利用Frank-Wolfe优化,在理论上达到极小极大下界,实验中将代理策略遗憾降低29.1%。
AI 中文摘要
从人类偏好中学习是大语言模型(LLM)对齐的核心,但人类偏好标注成本高昂。主动偏好学习通过选择信息量高的比较来降低成本,而LLM评判者可以提供额外的可扩展反馈。然而,评判者的偏好可能与目标人群的偏好存在偏差。即使在可信参考数据上校准后,主动获取也可能改变比较分布并暴露残余的评判者偏差。因此,我们将评判者偏差纳入获取设计,而不是依赖单独的校准阶段。在联合估计下,看似对奖励信息量高的比较也可能反映评判者偏差,从而提供较少关于人类偏好的信息。为解决此问题,我们提出了扰动调整最优设计(NAOD),一种比较选择策略,在扰动调整后优先考虑与策略相关的目标信息,并使用Frank-Wolfe算法进行优化。理论上,我们建立了策略风险上的尖锐条件局部渐近极小极大下界,并构造了达到该下界的估计器。我们进一步刻画了学习扰动表示的有限样本代价,并表明表示误差可以逆转最优设计的优势。最后,我们在实验上验证了这些预测,并在Chatbot Arena数据上评估NAOD,涵盖17个评判者、15种预算配置和15个随机聚类级划分。NAOD将代理策略的平均遗憾相对于匹配的目标信息设计降低了29.1%,优于现有方法,并改善了在留出数据上的人类偏好预测。
英文摘要
Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore provide less information about human preferences. To address this issue, we propose Nuisance-Adjusted Optimal Design (NAOD), a comparison-selection strategy that prioritizes policy-relevant target information after nuisance adjustment and uses the Frank-Wolfe algorithm for optimization. Theoretically, we establish a sharp conditional local asymptotic minimax lower bound on policy risk and construct an estimator that attains it. We further characterize the finite-sample cost of learning the nuisance representation and show that representation error can reverse an oracle design advantage. Finally, we validate these predictions experimentally and evaluate NAOD on Chatbot Arena data across 17 judges, 15 budget configurations, and 15 random cluster-level splits. NAOD reduces the mean regret of proxy policy by 29.1% relative to a matched target-information design, outperforms existing methods, and improves human-preference prediction on held-out data.
Comments35 pages, 5 figures