发表机构
NVIDIA Corporation; University of Michigan(英伟达公司; 密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出ADEPT强化学习框架,通过预训练和后训练实现高自由度机器人的可仿真到现实迁移的灵巧性,在两款机器人本体上完成长程任务,达到人类水平灵巧速度。
AI 中文摘要
我们提出了利用预训练提升灵巧性(Accelerating Dexterity via Pre-Training,简称ADEPT),这是一个大规模强化学习(RL)框架,用于学习跨高自由度(DoF)机器人本体的可从仿真到现实迁移的灵巧性,该框架可直接从原始视觉-触觉感知解决长程任务。ADEPT在通用物体重摆放任务上预训练灵巧策略,随后利用该预训练行为作为先验对下游策略进行后训练。ADEPT可学习多手指机器人从头开始难以发现的新行为,并避免为每个新下游任务重复学习相同技能集。预训练策略可零样本完成下游任务的重摆放阶段,但单纯的RL微调会在迁移过程中快速降低该能力。我们通过结合行为克隆蒸馏、评论者预热和保守在线更新的稳定后训练方案解决了该问题。为安全利用全运动学灵巧性,我们引入了关节空间几何结构(Geometric Fabric),用于在RL策略与机器人之间进行协调。我们将后训练的教师模型蒸馏为感知学生模型,使其在两个本体上实现从仿真到现实的零样本迁移:配备两个RGB相机的23 DoF Kuka-Allegro机器人,以及配备两个RGB相机和五个基于视觉的触觉传感器的29 DoF Flexiv-Sharpa机器人,该模型可从具有挑战性的初始状态以人类水平速度的灵巧性解决长程任务。
英文摘要
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real transferable dexterity across high degree-of-freedom (DoF) robot embodiments that can solve long-horizon tasks directly from raw visuo-tactile perception. ADEPT pretrains a dexterous policy on a generic object reposing task, then post-trains downstream policies with this pretrained behavior as a prior. ADEPT enables learning new behaviors that are otherwise difficult to discover from scratch on multi-fingered robots and avoids learning the same set of skills over again for every new downstream task. The pretrained policy zero-shots the reposing phase of downstream tasks, but naïve RL fine-tuning rapidly degrades this capability during transfer. We address this with a stable post-training recipe combining behavior-cloning distillation, critic warm-up, and conservative on-policy updates. To safely exploit the full kinematic dexterity, we introduce a joint-space Geometric Fabric that mediates between the RL policy and the robot. We distill post-trained teachers into perceptive students that zero-shot sim-to-real transfer on two embodiments: a 23 DoF Kuka-Allegro with two RGB cameras, and a 29 DoF Flexiv-Sharpa with two RGB cameras and five vision-based tactile sensors, and can solve long-horizon tasks from challenging initial states with dexterity at human-level speed.
CommentsProject page: https://adept-dexterity.github.io/