arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12651cs.LGcs.GT

重新审视AI对齐的失真:RLHF是一个体面的功利主义对齐器

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

  • University of California, Berkeley(加州大学伯克利分校)
  • Center for Advanced Intelligence Project, RIKEN(RIKEN先进智能项目中心)
  • The Institute of Statistical Mathematics and the Graduate University for Advanced Studies(统计数理研究所与综合研究大学院大学)
  • Tohoku University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

Kazusato Oko, Annie Ulichney, Nika Haghtalab, Han Bao

AI总结:

本文重新分析RLHF的失真问题,证明其指数退化源于分布不匹配而非算法缺陷,并给出紧界,表明无失配时RLHF可达最优失真。

AI中文摘要:

虽然基于人类反馈的强化学习(RLHF)是对齐大型语言模型与人类偏好的标准范式,但其在多元环境中的有效性已受到质疑。值得注意的是,Gölz等人(2025)的近期工作表明,当用户偏好异构时,失真(定义为RLHF策略的平均用户效用与最优平均效用之间的乘法差距)可能随Bradley-Terry温度参数β呈指数增长。在本工作中,我们对带奖励裁剪的RLHF失真进行了细粒度分析,并证明这种指数退化并非算法的固有属性,而是偏好数据生成分布(μ)与KL参考策略(π_ref)之间分布不匹配的结果。为此,我们在KL正则化强度的多个区间内建立了RLHF失真的紧上下界。我们表明,在代表性区间内,在Bradley-Terry模型下,失真为Θ̃(βB + β),其中B是μ与π_ref之间对数密度比的上界。特别地,当不存在分布不匹配(即μ = π_ref)时,RLHF在常数因子内达到最优失真O(β)。我们的结果表明,为了用RLHF合理最大化平均效用,最好使用策略内采样的偏好数据,或在RLHF之前对来自接近μ来源的数据进行微调。

英文摘要:

While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Gölz et al. (2025) demonstrated that the \textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $β$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a consequence of distribution mismatch between the distribution generating preference data ($μ$) and the KL reference policy ($π_{\mathrm{ref}}$). To this end, we establish tight upper and lower bounds on the distortion of RLHF across multiple regimes of the KL regularization strength. We show that in a representative regime, under the Bradley-Terry model, the distortion is $\tildeΘ(βB + β)$, where $B$ is an upper bound on the log density ratio between $μ$ and $π_{\mathrm{ref}}$. In particular, when there is no distribution mismatch (i.e., $μ= π_{\mathrm{ref}}$), RLHF achieves the optimal distortion of $O(β)$ up to a constant. Our results suggest that, to reasonably maximize average utility with RLHF, it is preferable to use on-policy sampled preference data or to fine-tune before RLHF on data from a source close to $μ$.

补充信息

↑