RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
RM-Distiller: 利用生成式大语言模型进行奖励模型蒸馏
机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) ; School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
AI总结 RM-Distiller通过系统利用教师LLM的精炼、评分和生成能力,提升奖励模型蒸馏效果,首次系统性地探索了生成式LLM在奖励建模中的应用。
Comments Accepted to IJCAI-ECAI 2026