Robo-Dopamine 2.0:面向机器人操纵的历史条件感知与分布外感知过程奖励建模
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
本文提出Robo-Dopamine 2.0,一种历史条件与分布外感知的过程奖励模型,结合带符号跳课程训练,提升了机器人操纵的视觉顺序一致性与分布外鲁棒性,在下游强化学习中取得优异的真实世界任务表现。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型可改进机器人操纵,但仍易受复合误差、场景变化和轨迹外状态影响。强化学习可优化预训练VLA策略,但稀疏成功信号阻碍探索,而人工设计的密集奖励成本高且任务特定。现有学习型视觉奖励模型常依赖静态前后观测,导致时间模糊性,且在分布外(OOD)执行时难以区分鲁棒性保留的变化与任务无效的失败。本文提出Robo-Dopamine 2.0,这是一种具有成对预测接口的历史条件与分布外感知过程奖励模型,它结合:(1)历史条件成对奖励,使用源对齐参考面板生成合成分布外查询,用观测的轨迹历史生成在线查询,同时保留查询端点;(2)分布外感知的带符号进展空间,表征有效进展、鲁棒性、失败和恢复。带转换感知回放的带符号跳课程先学习粗略执行顺序,再进行细粒度进展校准。本文还构建了分布外轨迹数据集和五组基准。参考面板将平均视觉顺序一致性(VOC)从0.967提升至0.986,分布外鲁棒VOC从0.906提升至0.958。在相同40万成对奖励预算下,采用25%回放的带符号跳训练达到0.9872的平均VOC,而配对池打乱对照为0.9858。在下游强化学习中,完整模型实现86.8%的平均RoboTwin成功率和80次真实世界插入任务中的71次成功。
英文摘要
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajectory states. Reinforcement learning can refine pretrained VLA policies, yet sparse success signals hinder exploration, while engineered dense rewards are costly and task-specific. Existing learned visual reward models often rely on static before-after observations, causing temporal ambiguity and weak discrimination between robustness-preserving variations and task-invalid failures under out-of-distribution (OOD) execution. We introduce Robo-Dopamine 2.0, a history- and OOD-aware process reward model with a pairwise prediction interface. It combines (1) history-conditioned pairwise rewards that use source-aligned reference panels for synthetic OOD queries and observed rollout history for online queries, while preserving the queried endpoints, and (2) an OOD-aware signed progress space that represents valid progress, robustness, failure, and recovery. A Signed-Hop Curriculum with transition-aware replay learns coarse execution ordering before fine-grained progress calibration. We also construct an OOD trajectory dataset and a five-family benchmark. Reference panels improve mean visual order consistency (VOC) from 0.967 to 0.986 and OOD-robust VOC from 0.906 to 0.958. With the same 400K pairwise-reward budget, Signed-Hop training with 25% replay reaches 0.9872 mean VOC, compared with 0.9858 for a matched-pool shuffled control. In downstream reinforcement learning, the full model achieves 86.8% mean RoboTwin success and 71/80 successful real-world insertions.
发表机构
- Peterson Co(彼得森公司)
机构由 AI 辅助整理,请以论文原文为准。