IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
IRPM:基于组间相对偏好的点wise生成奖励模型
机构 * HUJING Digital Media \& Entertainment Group (XingYun Lab), Beijing, China ; Department of Automation, Tsinghua University, Beijing, China ; Fudan University, Shanghai, China ; Beihang University, Beijing, China ; Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
AI总结 IRPM通过组间比较扩展Bradley-Terry模型,从成对偏好数据中训练点wise GRMs,实现高效可扩展的奖励建模。
Comments Comments: Updated title for clarity; improved theoretical derivations; added experiments at additional parameter scales and more ablations; added experimental details in the appendix; updated author list (added five co-authors) to reflect contributions to experiments and writing