Procedural Fairness Failures in RLHF from Preference Averaging
来自偏好平均的RLHF中的程序公平性失败
机构 * Vishnu Institute of Technology(维什努理工学院)
AI总结 该研究指出标准RLHF因偏好平均引发程序公平性失败,提出PA-RLHF分开优化不同偏好模式,提升了对齐准确率并缩小了群体公平差距,对大模型和智能体系统有重要意义。
Comments 4 pages, Accepted at the ICLR 2026 Workshop on Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA)