发表机构
Beihang University; Alibaba Group; Shanghai Innovation Institute; Sichuan University(北京航空航天大学; 阿里巴巴集团; 上海创新研究院; 四川大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EgoAlign通过执行反馈和控制器在环细化,将自我中心人类演示转化为适配人形机器人的动作监督,实现零样本远程操作-操控策略,提升物理拾取成功率并减少采集时间。
AI 中文摘要
以自我为中心的人类演示为任务经验提供了易于获取的来源,但身体比例和控制器响应的差异,以及机器人状态的缺失,限制了其作为人形机器人训练监督的价值。我们提出了EgoAlign,一个数据构建框架,将这些演示转换为与通用、连续的全身控制器兼容的动作和状态监督,而无需收集物理机器人演示。利用目标机器人模型和模拟器,EgoAlign通过执行反馈指导演示收集。它保留用于视觉引导的周期性步行的运动参考,同时通过尺度对齐和控制器在环细化来调整上半身交互几何。最终,因果回放重建相应的机器人状态和运动令牌标签,用于与人类观察一起训练。我们通过仅对适配的人类演示进行微调视觉-语言-动作模型,并在物理人形机器人上零样本部署来评估所得监督。所得策略执行远程物体重新定位、导航到未见过的目标位置以及独立评估的脚部交互。细化改善了模拟手部对齐和物理拾取成功率,相对于仅运动学对齐,而人类收集相对于遥操作减少了现场采集时间。此https URL
英文摘要
Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion references for visually guided periodic stepping while adapting upper-body interaction geometry through scale alignment and controller-in-the-loop refinement. A final causal replay reconstructs the corresponding robot states and motion-token labels for training with the human observations. We assess the resulting supervision by fine-tuning a vision--language--action model solely on adapted human demonstrations and deploying it zero-shot on a physical humanoid. The resulting policies perform long-range object relocation, navigation to unseen goal positions, and independently evaluated foot interaction. Refinement improves simulated hand alignment and physical pickup success over kinematic alignment alone, while human collection reduces on-site acquisition time relative to teleoperation. https://lambdahumanoid.github.io/EgoAlign/