OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
OpenRubrics: 向奖励建模和大语言模型对齐的可扩展合成 rubric 生成迈进
机构 * Purdue University(普渡大学) ; Emory University(埃默里大学) ; Georgia Institute of Technology(佐治亚理工学院) ; University at Albany(阿尔巴尼大学)
AI总结 OpenRubrics 提出了一种基于对比的 rubric 生成方法,通过大规模(提示,rubric)对提升奖励建模和大语言模型对齐的性能。
Comments The first two authors contributed equally. Updated OpenRubrics dataset, RMs, and results