arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15680cs.ROcs.AI

Robo-Dopamine 2.0:面向机器人操纵的历史条件感知与分布外感知过程奖励建模

Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation

Yijie Xu, Haopeng Jin, Run Zhou, Shengbang Liu, Sixiang Chen, Hongyang Cheng, Sicheng Hu, Peterson Co, Jinwen Luo, Huajie Tan, Shanghang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Robo-Dopamine 2.0,一种历史条件与分布外感知的过程奖励模型,结合带符号跳课程训练,提升了机器人操纵的视觉顺序一致性与分布外鲁棒性,在下游强化学习中取得优异的真实世界任务表现。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型可改进机器人操纵,但仍易受复合误差、场景变化和轨迹外状态影响。强化学习可优化预训练VLA策略,但稀疏成功信号阻碍探索,而人工设计的密集奖励成本高且任务特定。现有学习型视觉奖励模型常依赖静态前后观测,导致时间模糊性,且在分布外(OOD)执行时难以区分鲁棒性保留的变化与任务无效的失败。本文提出Robo-Dopamine 2.0,这是一种具有成对预测接口的历史条件与分布外感知过程奖励模型,它结合:(1)历史条件成对奖励,使用源对齐参考面板生成合成分布外查询,用观测的轨迹历史生成在线查询,同时保留查询端点;(2)分布外感知的带符号进展空间,表征有效进展、鲁棒性、失败和恢复。带转换感知回放的带符号跳课程先学习粗略执行顺序,再进行细粒度进展校准。本文还构建了分布外轨迹数据集和五组基准。参考面板将平均视觉顺序一致性(VOC)从0.967提升至0.986,分布外鲁棒VOC从0.906提升至0.958。在相同40万成对奖励预算下,采用25%回放的带符号跳训练达到0.9872的平均VOC,而配对池打乱对照为0.9858。在下游强化学习中,完整模型实现86.8%的平均RoboTwin成功率和80次真实世界插入任务中的71次成功。

英文摘要

Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.

发表机构

  • Peterson Co(彼得森公司)

机构由 AI 辅助整理,请以论文原文为准。

↑