arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2606.22027cs.ROcs.AI

RARM:基于置信度门控的进展奖励建模用于操作中的强化学习

RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

Pengzhi Yang, Xinyu Wang, Pengyu Jing, Kehan Wen, Yiduo Qu, Zhenhao Huang, Minghao Fu, Xin Liu, Yaheng Shen, Fan Shi

首次发表
浏览论文内容

中文总结 AI 辅助

提出参考锚定奖励模型(RARM),通过对比时间目标从通用视频中学习,将单个成功演示转化为密集进展奖励,并在部署时通过置信度门控抑制虚假正奖励,在长时域操作任务中显著提升强化学习成功率。

中文摘要 AI 辅助

机器人操作的强化学习常常受限于奖励设计,尤其是在长时域任务中:稀疏的成功奖励提供弱监督,而手工设计的密集奖励设计繁琐且跨任务泛化能力差。基于进展的奖励模型通过估计观察相对于任务完成的进展程度提供了一种有前景的替代方案,但现有方法通常需要任务特定的演示或进展标签,并且可能对视觉上合理但物理上错误的状态分配高奖励。我们引入了参考锚定奖励模型(RARM),一种轻量级视觉比较器,将单个成功演示转化为密集的、感知进展的奖励。RARM 在通用视频上使用对比时间目标训练一次,无需机器人特定数据、任务特定奖励标签或每任务奖励工程。在部署时,RARM 将 rollout 片段与参考片段匹配,并仅奖励置信的前向进展,抑制可能产生假阳性奖励的不确定匹配。在来自 LIBERO 和 MetaWorld 的 9 个模拟操作任务和 4 个真实世界任务中,RARM 在后续 RL 训练中取得了最佳总体成功率,在长时域任务(如布料折叠)中尤其有显著提升,因为在这些任务中不可靠的进展估计尤其有害。

英文摘要

Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poorly across tasks. Progress-based reward models offer a promising alternative by estimating how far an observation has advanced toward task completion, but existing approaches often require task-specific demonstrations or progress labels, and can assign high rewards to visually plausible but physically incorrect states. We introduce the Reference-Anchored Reward Model (RARM), a lightweight visual comparator that converts a single successful demonstration into a dense, progress-aware reward. RARM is trained once on general-purpose videos with a contrastive temporal objective, requiring no robot-specific data, task-specific reward labels, or per-task reward engineering. At deployment, RARM matches rollout clips to reference clips and rewards only confident forward progress, suppressing uncertain matches that may otherwise produce false-positive rewards. Across 9 simulated manipulation tasks from LIBERO and MetaWorld and 4 real-world tasks, RARM achieves the best overall success rates in subsequent RL training, with particularly large gains on long-horizon tasks such as cloth folding, where unreliable progress estimates are especially harmful.

发表机构

  • NUS Human-Centered Robotic Lab(新加坡国立大学人机共融机器人实验室)
  • University of Cambridge(剑桥大学)
  • School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

↑