无演示的成功概率奖励学习用于通用机器人策略
Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
查看机构详情
- Tsinghua University(清华大学)
- Yuanxing Robotics(元星机器人公司)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出无演示奖励学习范式,通过自举从稀疏任务结果学习成功概率,并引入RLER框架随策略进化调整奖励,显著提升机器人策略性能。
中文摘要 AI 辅助
强化学习(RL)使通用机器人策略能够通过试错交互进行改进,然而其有效性从根本上受到稀疏任务奖励的限制。现有的通用奖励模型通常通过从专家演示中学习任务进展来缓解这一问题,但引入了与策略优化过程中遇到的混合质量轨迹的分布不匹配,使得它们在策略必须学习的次优和失败行为上的估计不可靠。在这项工作中,我们提出了一种无演示的奖励学习范式,其中密集奖励反馈可以直接从稀疏任务结果和策略经验中学习。我们从理论上证明,终端任务结果隐式定义了中间时间步的密集成功概率反馈,可以通过自举递归学习。基于这一见解,我们引入了eVTA$_0$,它通过时间差分式自举从混合质量的策略轨迹中学习成功概率,无需专家演示或中间注释。我们进一步引入了具有进化奖励的强化学习(RLER),这是一个闭环框架,随着策略的进化,使用新收集的轨迹来适应eVTA$_0$。实验表明,在相同的RL训练预算下,eVTA$_0$提供了比最先进的奖励模型更具信息量的奖励,并在所有LIBERO任务套件中实现了最佳的平均策略性能,相对于初始策略将成功率提高了5.4%-13.8%。在真实世界的操作中,RLER进一步将总体成功率提高了20%-26%,在分布外条件下获得了35%-36%的提升。这些结果证明了无演示奖励学习和随着策略进化调整奖励的有效性。项目网页:此https URL。
英文摘要
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA$_0$, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA$_0$ using newly collected rollouts as the policy evolves. Experiments show that eVTA$_0$ provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.