来自偏好平均的RLHF中的程序公平性失败
Procedural Fairness Failures in RLHF from Preference Averaging
查看机构详情
- Vishnu Institute of Technology(维什努理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究指出标准RLHF因偏好平均引发程序公平性失败,提出PA-RLHF分开优化不同偏好模式,提升了对齐准确率并缩小了群体公平差距,对大模型和智能体系统有重要意义。
中文摘要 AI 辅助
人类反馈强化学习(RLHF)将异质偏好聚合为单一奖励模型,其前提是偏好具有同质性。当偏好异质时,这种聚合会引发程序公平性失败:多数偏好群体主导奖励学习,而少数群体的偏好则被系统性低估。本研究将对齐过程中的程序公平性定义为在奖励建模阶段保留不同的偏好信号,结果表明标准RLHF通过偏好平均违反了该定义。研究提出偏好感知RLHF(PA-RLHF),在奖励学习阶段对不同偏好模式分开优化。在受控环境中,PA-RLHF将整体对齐准确率从46.9%提升至67.9%,并将对齐最优与最差群体间的公平差距从15.9个百分点缩小至9.6个百分点。这些结果表明,即使在受控无噪声环境中,对齐过程中的程序公平性失败也可能源于奖励学习的结构设计选择,这对大型语言模型和智能体系统具有直接影响,因为有偏奖励模型会在序列决策中加剧不公平。
英文摘要
Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness failure where majority preference groups dominate reward learning while minority preferences are systematically under-represented. This work defines procedural fairness in alignment as preserving distinct preference signals during reward modeling and shows that standard RLHF violates this via preference averaging. Preference-Aware RLHF (PA-RLHF) is introduced, separating optimization across preference modes at the reward learning stage. In a controlled setting, PA-RLHF improves overall alignment accuracy from 46.9% to 67.9% and reduces the fairness gap between best and worst aligned groups from 15.9 to 9.6 percentage points. These results show that procedural fairness failures in alignment can arise from structural design choices in reward learning, even in controlled, noise-free settings, with direct implications for large language models and agentic systems, where biased reward models can compound inequities across sequential decisions.