发表机构
School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; JD Explore Academy(中国科学院大学人工智能学院; 中国科学院自动化研究所; 京东探索学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Dream2Reward从正演示学习成功潜在转换场,通过转换级比较生成密集因果奖励,可提升机器人操纵的成功-失败区分度与下游策略性能。
AI 中文摘要
学习机器人策略需要密集奖励,该奖励在行为偏离成功演示时仍需具备信息性。基于进度的奖励会估计观测值沿标称成功轨迹的推进程度,但可能在错误转换后仍保持较高值。我们提出Dream2Reward,它从正演示中学习语言条件下的成功潜在转换场。给定转换起始点的视觉历史,模型会预测与成功执行相关的潜在位移,并通过符号方向和对称幅度一致性对观测位移进行评分。这种转换级比较会惩罚错误方向、过冲和停滞运动,即使产生的观测值看似显示出进展。Dream2Reward无需失败标注、进度标签或合成负样本,即可生成密集的因果奖励。在机制诊断和共享轨迹评估中,它相比基于进度的替代方案,能提供更强的成功-失败区分度和对低质量行为更具信息性的反馈。在在线和离线策略学习中,同一冻结奖励模型可减少奖励黑客行为并支持更强的下游性能,包括在真实机器人操纵中的表现。这些结果表明,将实际运动与预测的成功变化进行比较,是将正演示转换为机器人学习密集奖励的有效方法。
英文摘要
Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations. Given the visual history up to a transition start, the model predicts the latent displacement associated with successful execution and scores the observed displacement through signed directional and symmetric magnitude agreement. This transition-level comparison penalizes wrong-direction, overshooting, and stagnant motion even when the resulting observation appears to show progress. Dream2Reward requires no failure annotations, progress labels, or synthetic negatives, and produces a dense causal reward. Across mechanism diagnostics and shared-trajectory evaluations, it provides stronger success-failure separation and more informative feedback on low-quality behavior than progress-based alternatives. Across online and offline policy learning, the same frozen reward model reduces reward hacking and supports stronger downstream performance, including in real-robot manipulation. These results show that comparing realized motion with predicted successful change provides an effective way to convert positive demonstrations into dense rewards for robot learning.
Comments12 pages, 7 figures