arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Open-AoE:用于具身学习的开放自我中心操纵数据集和工具链

Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Zishuo Li, Bowen Yang, Changtao Miao, Kai Zhu, Hao Chen, Qingze Guan, Zhengxing Wu, Wanke Zhan, Yang Sun, Zhiyi Huang, Zitong Shan, Zhenchao Jin, Jiadong Hong, Taowen Wang, Yushi Feng, You Liu, Yibo Wang, Yifan Yang, Zhaowen Zhou, Man Luo, Hao Cheng, Bo Zhang, Jianshu Li, Jiansheng Cai, Guocai Yao, Jize Zhang, Chenhao Lin, Renjing Xu, Lequan Yu, Chao Shen, Chunhua Shen, Zhe Li

arXiv 2607.14183首次发表:更新:

发表机构

Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在为具身学习提供数据集和工具链,提出Open-AoE。它涵盖从手机捕捉到模型训练全流程,含约2000小时视频及多种注释等。通过整合多环节,降低数据贡献与重用障碍,为相关研究提供实用开放基础设施。

AI 中文摘要

人类操纵的自我中心视频为具身智能提供了可扩展的监督,但现有资源很少能将低成本的连续捕捉、操纵级别的结构化注释以及用于机器人学习的可重复使用工具结合起来。我们提出了Open-AoE,这是一个开放的、面向社区的自我中心操纵数据集和工具链,涵盖了从智能手机捕捉到模型训练的完整流程。其首个版本包含由500多名贡献者使用400多部智能手机在自然环境中收集的约2000小时操纵视频。该数据集提供文本注释、基于MANO的手部姿势、相机轨迹以及时间定位的原子动作。Open-AoE还包括一个数据处理管道,通过时间动作分割、语义注释、手部重建和相机轨迹重建将原始记录转换为结构化样本。同时,我们提供了一个单独的下游工具链,支持可视化、跨具身重定向、特定模型的数据转换以及针对VLA策略、WAMs和世界模型的训练方法。通过整合可扩展的捕捉、结构化处理和下游适配,Open-AoE降低了数据贡献和重用的障碍,为具身模型训练、人机转移和世界建模提供了实用的开放基础设施。

英文摘要

Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑