Dexplore:基于参考范围探索的可扩展灵巧操作神经控制
Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration
浏览论文内容
中文总结 AI 辅助
Dexplore提出统一单循环优化,直接从运动捕捉数据联合重定向与跟踪,以软引导和自适应范围训练策略,蒸馏为视觉技能控制器,实现可扩展的灵巧操作。
中文摘要 AI 辅助
手-物体运动捕捉(MoCap)数据库提供了大规模、富含接触的演示,有望用于扩展灵巧机器人操作。然而,演示的不准确性以及人手与机器人手之间的具身差异限制了这些数据的直接使用。现有方法采用三阶段工作流程,包括重定向、跟踪和残差校正,这往往导致演示未被充分利用,并在各阶段之间累积误差。我们提出了Dexplore,一种统一的单循环优化方法,联合执行重定向和跟踪,直接从MoCap大规模学习机器人控制策略。我们不将演示视为真值,而是将其用作软引导。从原始轨迹中,我们推导出自适应空间范围,并使用强化学习进行训练,使策略保持在范围内,同时最小化控制努力并完成任务。这种统一公式保留了演示意图,使机器人特定策略得以涌现,提高了对噪声的鲁棒性,并可扩展到大型演示语料库。我们将扩展的跟踪策略蒸馏为基于视觉的、技能条件的生成控制器,该控制器在丰富的潜在表示中编码多种操作技能,支持跨对象的泛化和现实世界部署。综合来看,这些贡献使Dexplore成为一座原则性的桥梁,将不完美的演示转化为灵巧操作的有效训练信号。
英文摘要
Hand-object motion-capture (MoCap) repositories offer large-scale, contact-rich demonstrations and hold promise for scaling dexterous robotic manipulation. Yet demonstration inaccuracies and embodiment gaps between human and robot hands limit the straightforward use of these data. Existing methods adopt a three-stage workflow, including retargeting, tracking, and residual correction, which often leaves demonstrations underused and compound errors across stages. We introduce Dexplore, a unified single-loop optimization that jointly performs retargeting and tracking to learn robot control policies directly from MoCap at scale. Rather than treating demonstrations as ground truth, we use them as soft guidance. From raw trajectories, we derive adaptive spatial scopes, and train with reinforcement learning to keep the policy in-scope while minimizing control effort and accomplishing the task. This unified formulation preserves demonstration intent, enables robot-specific strategies to emerge, improves robustness to noise, and scales to large demonstration corpora. We distill the scaled tracking policy into a vision-based, skill-conditioned generative controller that encodes diverse manipulation skills in a rich latent representation, supporting generalization across objects and real-world deployment. Taken together, these contributions position Dexplore as a principled bridge that transforms imperfect demonstrations into effective training signals for dexterous manipulation.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。