arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IronMind:通过相机空间自我中心预训练扩展人形机器人灵巧操作能力

IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining

Huimin Pan, Yufan Ren, Kunpeng Song, Siyang Wang, Xiwen Zhang, Xiaoyun Hu, Zhuoxu Duan, Hanrui Zheng, Jialeng Ni, Nathan Zhao, Sibo Ma, Zhenxuan Fan, Zhongyang Che, Danny Bao, Jiacheng Wei, Jerry Bai, Xiaoyu Yue, Xiaoyang Guo, Chenyi Chen

arXiv 2609.39403首次发表:更新:

发表机构

XPENG Robotics(小鹏机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IronMind通过相机空间动作表示,利用大规模人类自我中心视频预训练VLA模型,弥合具身差异,显著提升人形机器人灵巧操作的成功率(10,000小时预训练达55.0%)。

AI 中文摘要

自我中心的人类视频为灵巧操作提供了可扩展的数据源,但利用这些数据训练人形机器人面临两大挑战:(1)具身差异,即人类手部与机器人末端执行器在结构上存在差异,且低成本自我中心记录缺乏传统重定向所需的躯干运动学信息;(2)异构数据质量,包括嘈杂的手部姿态跟踪和弱对齐的文本标注。我们提出了IronMind,一种视觉-语言-动作(VLA)模型,利用自我中心人类视频和异构机器人数据来预训练人形机器人灵巧操作策略。为弥合具身差异,IronMind通过使用相机空间动作表示(自我中心视频的原生参考空间)并语义对齐机器人与人类动作维度,绕过了显式的身体重定向。在从250到10,000小时的总预训练预算范围内,验证损失随数据规模近似对数线性下降。更大的预训练预算还改善了后训练后的分布外真实机器人操作:在涉及未见物体、未见功能属性和推理提示的六项具有挑战性的任务中,10,000小时模型实现了55.0%的成功率,而所有不超过5,000小时的预训练预算的成功率最高仅为11.7%,无预训练时则为5.0%。在相同的预训练预算下,相机空间动作表示也优于躯干框架基线。综合这些发现,支持将大规模人类自我中心数据上的相机空间动作表示预训练作为人形机器人操作的可扩展基础。

英文摘要

Egocentric human video offers a scalable data source for dexterous manipulation, yet using it to train humanoid robots presents two challenges: (1) an embodiment gap, as human hands differ structurally from robot end-effectors and low-cost egocentric recordings lack the torso kinematics required by conventional retargeting; and (2) heterogeneous data quality, including noisy hand-pose tracking and weakly aligned text annotations. We introduce IronMind, a vision-language-action (VLA) model that uses egocentric human video and heterogeneous robot data to pretrain policies for humanoid dexterous manipulation. To bridge the embodiment gap, IronMind bypasses explicit body-retargeting by using a camera-space action representation, the native reference space of egocentric video, and semantically aligning robot and human action dimensions. Across total pretraining budgets from 250 to 10,000 hours, validation loss decreases approximately log-linearly with data scale. Larger pretraining budgets also improve out-of-distribution real-robot manipulation after post-training: across six challenging tasks with unseen objects, affordances, and reasoning prompts, the 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining. At the same pretraining budget, the camera-space action representation also outperforms the torso-frame baseline. Together, these findings support pretraining with a camera-space action representation on large-scale human egocentric data as a scalable foundation for humanoid robot manipulation.

Commentshttps://xpeng-robotics.github.io/ironmind/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑