Real-Time Aligned Reward Model beyond Semantics
实时对齐奖励模型超越语义
机构 * Beihang University(北京航空航天大学) ; Tsinghua University(清华大学) ; Renmin University of China(中国人民大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
专题命中 偏好对齐 :RLHF(summary_cn,abstract);分类 cs.AI
AI总结 本文提出R2M框架,通过实时利用策略模型反馈来对齐策略分布偏移,解决RLHF中奖励过拟合问题。