Why is Your Language Model a Poor Implicit Reward Model?
为什么你的语言模型是一个差的隐式奖励模型?
机构 * Some Institute(某些研究所)
专题命中 后训练与偏好优化 :language model(title,abstract);post-training(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 研究发现隐式奖励模型(IM-RM)在泛化能力上劣于显式奖励模型(EX-RM),因其更依赖表面token级线索,而设计选择对模型泛化行为有显著影响。
Comments Accepted to ICLR 2026; Code available at https://github.com/princeton-pli/exrm-vs-imrm