AtomEgo:探索具身基础模型预训练中的自我-机器人集成
AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining
浏览论文内容
中文总结 AI 辅助
AtomEgo系统研究自我-机器人协同训练,利用2659小时第一人称数据,提出数据规模与对齐质量决定能力增益的原则,指导具身基础模型预训练。
中文摘要 AI 辅助
具身基础模型受限于机器人演示数据的规模和多样性,这促使我们利用大规模的第一人称视角人类交互数据。然而,由于人类与机器人在具身形态和动作空间上存在显著差异,如何有效地将这些数据整合到具身模型预训练中仍不明确。我们提出了AtomEgo,一项关于自我-机器人协同训练的系统性研究,该研究依托一个约2659小时的精选语料库和可扩展的数据处理流水线。在视觉-语言-动作和世界-动作模型架构中,我们研究了三种代表性范式:使用特定领域动作头的联合协同训练、通过具身对齐进行的渐进式自我到机器人迁移,以及联合视频-动作建模。我们通过多任务真实机器人实验和语言条件跨具身表征分析来评估这些范式。我们的结果揭示了一个简单原则:数据规模 * 对齐质量 --> 能力增益;第一人称视角数据可以提高泛化能力,但其价值取决于它们被对齐和利用的有效程度。这一原则可以为可扩展的自我-机器人预训练提供实用指导。
英文摘要
Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision--language--action and world--action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video--action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality --> Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego--robot pre-training.
发表机构
- INFIFORCE
- Lionrock Artificial Intelligence Laboratory(狮岩人工智能实验室)
- HKU(香港大学)
- Tongji University(同济大学)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。