Revisiting the Learning Objectives of Vision-Language Reward Models
重新审视视觉-语言奖励模型的学习目标
机构 * Polytechnique Montréal(蒙特利尔理工学院) ; École de Technologie Supérieure(高级技术学院) ; Sorbonne Université(索邦大学)
专题命中 具身与机器人 :world model(comments);分类 cs.AI、cs.LG;world model(comments)
AI总结 本文通过统一框架评估了基于VLM的奖励模型,发现简单三元组损失在性能上优于现有方法,表明改进可能源于数据和架构差异。
Comments Published as an extended abstract at World Modeling Workshop 2026