Think Twice: Branch-and-Rethink Reasoning Reward Model
再思两次:分支与再思考推理奖励模型
机构 * NVIDIA(NVIDIA公司) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
AI总结 BR-RM通过两轮次的分支与再思考机制,改进奖励模型的判断准确性与敏感性,实现更精确的推理能力。
Comments Source Code: https://github.com/yzjiao/BR-RM. Model Checkpoints: https://huggingface.co/nvidia/Qwen3-Nemotron-14B-BRRM and https://huggingface.co/nvidia/Qwen3-Nemotron-8B-BRRM