发表机构
The Hong Kong University of Science and Technology; Beijing Academy of Artificial Intelligence; State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; Beihang University; Institute of Automation, Chinese Academy of Sciences(香港科技大学; 北京人工智能研究院; 北京大学计算机学院多媒体信息处理国家重点实验室; 北京航空航天大学; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DexRoam提出无追踪器采集系统与三阶段对齐方法,利用自我中心全身人体演示提升移动双臂灵巧操作策略学习,将成功率提升至约56-57%。
AI 中文摘要
移动双臂灵巧操作要求在单一轨迹中持续协调运动、全身运动与手指级灵巧性,这造成了严重的机器人演示瓶颈。自我中心人体演示提供了一种可扩展的替代方案,但先前的方法通过简化人体运动来缓解迁移难度,恰恰丢弃了此类任务所依赖的细粒度耦合结构。我们提出DexRoam,一个从人体演示中学习移动双臂灵巧操作的完整系统,在整个从人到机器人的迁移过程中,全身运动保持连续且耦合。为了实现可扩展的全身人体操作演示收集,我们开发了一种无追踪器的采集系统,仅使用消费级VR头显和头戴式立体相机,无需外部相机或运动追踪器。随后,我们执行三个明确的对齐阶段——具身对齐、动作语义对齐和时间对齐——将捕获的运动映射到机器人动作空间,保留细粒度的全身运动,并允许人体与机器人演示由标准VLA策略联合学习。使用不同VLA主干网络的真实世界实验表明,人体演示在不同训练范式下持续改善策略学习,在GR00T N1.7上将平均成功率从29%提升至56%,在pi0.5上从32%提升至57%,同时使用一半的机器人演示即可匹配仅用机器人训练的效果。消融实验确认每个对齐阶段都是必要的。这些结果凸显了人体演示在保留细粒度运动结构的可扩展全身移动操作中的潜力。
英文摘要
Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease the transfer by simplifying human motion, discarding exactly the fine-grained, coupled structure such tasks depend on. We present DexRoam, a complete system for learning mobile bimanual dexterous manipulation from human demonstrations, in which whole-body motion remains continuous and coupled throughout the human-to-robot transfer process. To enable scalable collection of whole-body human manipulation demonstrations, we develop a tracker-free capture system using only a consumer VR headset and a head-mounted stereo camera, without external cameras or motion trackers. We then perform three explicit alignment stages---embodiment, action-semantic, and temporal---to map captured motion into the robot action space, preserving fine-grained whole-body motion and allowing human and robot demonstrations to be jointly learned by standard VLA policies. Real-world experiments with different VLA backbones show that human demonstrations consistently improve policy learning across training paradigms, raising average success from 29% to 56% on GR00T N1.7 and from 32% to 57% on pi0.5, while matching robot-only training with half the robot demonstrations. Ablations confirm that each alignment stage is necessary. These results highlight the potential of human demonstrations for scalable whole-body mobile manipulation with preserved fine-grained motion structure.
CommentsProject page: https://dexroam.github.io/