发表机构
Lipscomb University(利普斯科姆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SymmGrid是一种并行对称性轨迹级增强框架,通过自我-外部视觉感知加速机器人上学习,在真实机器人操作任务上实现训练速度、成功率及nAUC比率提升,最快收敛时间低至10.9分钟。
AI 中文摘要
深度强化策略学习在物理机器人上直接开展(机器人上学习)仍受限于实际训练时间过长的瓶颈。我们提出SymmGrid,这是一种受并行对称性启发的轨迹级增强框架,可对群变换进行超规模扩展,从而在自我感知和外部感知视觉设置中显著加速机器人上学习。我们在对称树下对马尔可夫决策过程(MDP)进行建模,其中状态-动作对具有可允许的并行不变变换,这些变换会生成几何网格结构。状态由自我或外部图像以及本体感觉信息建模,后者需要通过单应性(homographies)进行特殊处理,以根据对应的空间变换对视觉场景进行扭曲。这些并行变换会产生大量独特的对称等价关系,在回放缓冲区中填充多样化且一致的经验,从而加速学习并提升性能。我们在真实机器人操作接触任务上开展了广泛的训练与评估,包括插销、布线和物体移位。相较于SOTA,SymmGrid实现了1.37至2.17倍的实际训练收敛速度提升、1.09至1.27倍的评估成功率提升,各任务最快训练收敛时间分别为16.6分钟、10.9分钟和79.3分钟。对于轨迹范围评估,我们使用了归一化曲线下面积(nAUC)比率,SymmGrid实现了最高2.59倍的提升。这些结果证实,简单的分支对称性可因超规模扩展产生显著效果,让我们更接近实现机械臂和人形机器人适用的操作任务中10分钟以内的机器人上学习训练。项目页面可通过该http URL访问。
英文摘要
Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times. We present SymmGrid, a trajectory level augmentation framework inspired by parallelized symmetries that super-scales group transformations to significantly accelerate on-robot learning in both egocentric and exocentric visual setups. We model a Markov Decision Process (MDP) under a symmetry tree, in which state-action pairs have admissible parallelized invariant transformations that yield a geometric grid structure. The state is modelled with ego- or exocentric images and proprioception information. The latter require special treatment, in the form of homographies, to warp visual scenes in line with their corresponding spatial transformations. These parallelized transformations produce a large set of unique symmetric equivalences that populate the replay buffer with diverse and consistent experiences that speed up learning and improve performance. We present extensive training and evaluations performed directly on real robot manipulation contact tasks including peg-insertions, cable routing, and object relocations. Relative to SOTA, SymmGrid achieved wall-clock training convergence speed-ups of 1.37-2.17x, evaluation success rate improvements of 1.09x-1.27x, fastest training convergence times of 16.6, 10.9, and 79.3 minutes respectively. For trajectory wide assessments, we used normalized area under the curve (nAUC) ratios. SymmGrid achieved improvements of up to 2.59x. These results confirm that simple branch symmetries can have an outsized result due to super-scaling and bring us closer to sub-10 minute on-robot learning training in manipulation tasks suitable for arms and humanoids. The project page is available at symmgrid-robot.github.io
Comments9 pages, 7 figures, 1 table