Process Reward Models That Think
能够思考的过程奖励模型
机构 * University of Michigan(密歇根大学) ; Mila ; LG AI Research(LG人工智能研究) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Equal supervision(共同监督)
AI总结 ThinkPRM通过生成长CoT验证链实现高效过程奖励模型,以更少的监督标签在多个基准测试中超越现有方法。
Comments Add new ablation and minor writing fixes