arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21788cs.ROcs.LG

从预训练到精通:基于真实世界子任务强化学习的最小人工干预长时程操作

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

  • UT Austin(德克萨斯大学奥斯汀分校)
  • Autel US(Autel美国)
  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun

AI总结:

提出PARTS框架,通过聚焦关键子任务的真实世界强化学习,以最少人工干预提升长时程操作成功率,在双臂和单臂任务上分别提升至61%和95%。

AI中文摘要:

预训练的机器人基础策略可以执行长时程任务的大部分步骤,但会在少数关键子任务上反复失败。为监督微调(SFT)收集额外的完整任务演示,要求操作员重复策略已经表现良好的行为。强化学习(RL)微调为弥合这一差距提供了一条有前景的路径,但现有方法难以仅使用稀疏奖励解决长时程任务。我们提出了PARTS(基于目标子任务强化学习的策略自适应),一个真实世界的子任务强化学习框架,将练习集中在这些瓶颈上,同时允许训练回滚以最少的人工干预进行。冻结的预训练策略在整个执行过程中提供名义动作,而智能体生成的选择器和成功验证器激活残差修正并提供局部结果奖励。这些奖励支持在完整任务成功稀少时从成功的子任务中学习。训练将在线强化学习与成功重新加权的重训练相结合,每个重训练的残差策略被重新部署以收集更多经验。人类在设置期间识别瓶颈,并在需要时进行物理重置。在双臂YAM和单臂Franka任务上,PARTS将完整任务成功率分别从32%提高到61%,从50%提高到95%,每个任务平均使用数十分钟的真实世界强化学习回滚。与现有的真实世界强化学习微调方法相比,PARTS在相同的机器人回滚预算下将完整任务成功率提高了超过25%,同时需要更少的人工参与。

英文摘要:

A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.

补充信息

↑