发表机构
ETH Zürich; NVIDIA(苏黎世联邦理工学院; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对稀疏奖励长时程任务,提出双向Voronoi偏置探索课程(BVER),从目标与初始状态两端同时扩展,加速目标条件强化学习,实验显示显著优于单侧扩展方法。
AI 中文摘要
具有稀疏奖励的长时程任务对目标条件强化学习构成了探索瓶颈:从初始状态出发的策略很少能到达目标,因而无法获得学习信号。参考运动、手工设计的课程和塑形奖励可以提供这种信号,但需要演示或特定任务的工程;自动的起始状态和目标课程避免了这些需求,但通常仅从一侧扩展,因此必须从该侧覆盖到目标的全部距离。我们提出了双向Voronoi偏置探索课程用于强化学习(BVER),该方法同时从两端扩展。受双向RRT规划启发,BVER从目标向外生长起始状态,并从初始状态分布向外生长目标,将两者都偏向未探索的任务空间,并引导它们相互靠近,训练一个目标条件策略同时处理两者。在质点迷宫、四足机器人爬箱和机械臂环上杆转移任务中,BVER比所有比较的无参考课程学习得更快。在爬箱任务中,对于0.4米高的箱子,BVER达到95%成功率所需的迭代次数比最佳比较方法少约65%,并且是唯一能学会爬0.7米高箱子的方法,同时产生的策略对起始位置、目标位置和偏航角变化具有鲁棒性。在没有演示的情况下,BVER在0.4米箱子任务和环上杆转移任务上接近基于参考的课程的样本效率。消融实验表明,从两端扩展优于仅从任一方向扩展。
英文摘要
Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.