arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DexWeave:从人类演示中学习灵巧的人形移动操作

DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations

Naichuan Sun, Haotian Shen, Yizhang Zhang, Luying Feng, Haoze Wang, Yuanbo Xiangli, Yaochu Jin, Peidong Liu

arXiv 2609.34724首次发表:更新:

发表机构

Westlake University; Shanghai Jiao Tong University(西湖大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DexWeave提出统一框架,通过交互一致的运动重定向和解剖感知的Transformer策略,从人类演示学习灵巧人形移动操作,实现更高性能与快速收敛,并在真实机器人上验证。

AI 中文摘要

从人类演示中学习灵巧的人形移动操作,不仅需要迁移人类的运动,还需要迁移演示行为背后的协调交互结构。这具有挑战性,因为具身差异扭曲了身体运动、手腕放置、手指关节运动和物体交互之间的耦合,而运动学上准确的参考在机器人动力学下可能仍然难以实现。我们提出了DexWeave,一个统一框架,将交互一致的运动重定向与解剖感知的全身策略学习相结合。DexWeave首先采用两阶段重定向程序,用专门的求解器初始化身体和手部运动,随后在上半身交互链上进行耦合细化,同时保持下半身支撑。生成的参考由解剖感知的Transformer策略跟踪,该策略将解剖区域表示为结构化标记,并使用定向掩码注意力来建模它们的依赖关系,物体信息选择性地调节上半身路径以实现灵巧交互。该策略联合输出身体和灵巧手动作,并直接使用强化学习训练,无需预训练跟踪策略、教师-学生蒸馏或后续残差细化。DexWeave提高了重定向保真度和交互一致性,同时实现了比MLP基线更高的操作性能和更快的策略收敛。我们进一步将学习到的策略部署在配备Inspire灵巧手的物理Unitree G1人形机器人上,在现实世界中展示了灵巧的全身移动操作。视频请参见我们的项目页面(此https URL)。

英文摘要

Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist placement, finger articulation, and object interaction, while kinematically accurate references may still be difficult to realize under robot dynamics. We present DexWeave, a unified framework that connects interaction-consistent motion retargeting with anatomy-aware whole-body policy learning. DexWeave first employs a two-stage retargeting procedure that initializes body and hand motions with specialized solvers and subsequently performs coupled refinement over the upper-body interaction chain while preserving lower-body support. The resulting references are tracked by an anatomy-aware Transformer policy that represents anatomical regions as structured tokens and uses directed masked attention to model their dependencies, with object information selectively conditioning the upper-body pathway for dexterous interaction. The policy jointly outputs body and dexterous-hand actions and is trained directly with reinforcement learning, without pretrained tracking policies, teacher-student distillation, or subsequent residual refinement. DexWeave improves retargeting fidelity and interaction consistency while achieving higher manipulation performance and faster policy convergence than MLP baselines. We further deploy the learned policies on a physical Unitree G1 humanoid equipped with Inspire dexterous hands, demonstrating dexterous whole-body loco-manipulation in the real world. See our project page (https://dexweave.github.io) for videos.

Comments26 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑