Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai AI Lab(上海人工智能实验室) ; Hunyuan, Tencent(腾讯 Hunyuan)
专题命中 测试时计算 :chain-of-thought(title);reasoning(abstract);CoT(abstract)
Comments [NeurIPS2025] Project Page: https://codegoat24.github.io/UnifiedReward/think