Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) ; Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) ; Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部) ; Kuaishou Technology(快手科技) ; University of Chinese Academy of Sciences(中国科学院大学)
Comments Accepted to ACL 2025