发表机构
Peking University; Beijing Innovation Center of Humanoid Robotics; University of Science and Technology of China; Southeast University; Southern University of Science and Technology; Beijing University of Aeronautics and Astronautics; Beijing Language and Culture University; Sichuan University; Beijing University of Posts and Telecommunications(北京大学; 北京人形机器人创新中心; 中国科学技术大学; 东南大学; 南方科技大学; 北京航空航天大学; 北京语言大学; 四川大学; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有机器人奖励模型的跨范式偏好与分数不一致问题,本文提出 TrustRoboReward 框架,通过 POISE 方法解决反转冲突,训练的 Qwen3-VL-4B 性能接近 GPT-5-mini,优于 RoboReward 基线。
AI 中文摘要
奖励模型是具身人工智能中强化学习的瓶颈。长 horizon 机器人操作需要超越人工设计奖励或任务特定标注的可扩展视觉反馈。现有开源 VLM 奖励评判器如 RoboReward 采用简单的 1-5 轨迹进度评分,缺乏 RLHF、DPO 和 Bradley-Terry 框架所需的成对偏好,且无法优化视频场景理解。为 RoboReward 添加成对比较和视频 QA 监督会导致成对偏好与点态分数不一致,引入训练噪声并损害下游性能——TrustJudge 等聚合方法无法解决该问题。为此,本文提出 TrustRoboReward,一种配备偏好有序保序分数编辑(POISE)的多范式奖励建模框架。我们构建了包含轨迹进度评分(Score-A)、视频 QA 答案质量评分(Score-B)及其成对对应项(Pair-A、Pair-B)的统一四范式数据集。成对标签比点态分数更符合人类判断,启发我们校准点态分数以避免与成对偏好的分数对反转。POISE 可纠正点态分数,消除 TrustJudge 无法解决的跨范式反转冲突。理论上,POISE 将分数对反转冲突从 20.15% 降至 0%,而 TrustJudge 在相同语料库上保留 20.46% 的冲突。在我们的基准上,经 POISE 训练的 Qwen3-VL-4B 获得 77.96% 的整体奖励分数,几乎与 GPT-5-mini(78.09%,差距 0.13%)持平,且比最强的 RoboReward-4B 基线高出 10.13%。它还将测试时分数对一致性提升至 71.90%,超过 RoboReward-4B(57.26%)和 GPT-5-mini(68.09%)。推理时集成 TrustJudge 聚合可将整体分数提升至 78.57%,超越 GPT-5-mini 教师模型。
英文摘要
Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedback beyond handcrafted rewards or task-specific annotations. Existing open-source VLM reward judges like RoboReward adopt simple 1--5 trajectory progress scoring, lacking pairwise preferences for RLHF, DPO and Bradley-Terry frameworks, while failing to optimize video scene understanding. Augmenting RoboReward with pairwise comparison and video-QA supervision causes inconsistency between pairwise preferences and pointwise scores, introducing training noise and hurting downstream performance---an issue aggregation methods such as TrustJudge cannot resolve. To address this, we propose TrustRoboReward, a multi-paradigm reward modeling framework equipped with Preference-Ordered Isotonic Score Editing (POISE). We construct a unified four-paradigm dataset with trajectory progress scoring (Score-A), video-QA answer quality scoring (Score-B), and their pairwise counterparts (Pair-A, Pair-B). Pairwise labels align better with human judgment than pointwise scores, inspiring us to calibrate pointwise scores to avoid score-pair reversals against pairwise preferences. POISE rectifies pointwise scores and eliminates cross-paradigm reversal conflicts unresolved by TrustJudge. Theoretically, POISE reduces score-pair reversal conflicts from 20.15% to 0%, whereas TrustJudge retains 20.46% conflicts on the same corpus. Evaluated on our benchmark, Qwen3-VL-4B trained with POISE achieves an overall reward score of 77.96%, nearly matching GPT-5-mini (78.09%, gap 0.13%) and outperforming the strongest RoboReward-4B baseline by 10.13%. It also lifts test-time score-pair consistency to 71.90%, exceeding RoboReward-4B (57.26%) and GPT-5-mini (68.09%). Integrating TrustJudge aggregation during inference boosts the overall score to 78.57%, surpassing the GPT-5-mini teacher model.