OpenEgo:面向灵巧操作的大规模多模态自我中心数据集
OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation
浏览论文内容
中文总结 AI 辅助
OpenEgo是一个大规模多模态自我中心操作数据集,含1107小时视频、290个任务,提供标准化手部姿态和动作原语,用于降低灵巧操作模仿学习门槛并支持可复现研究。
中文摘要 AI 辅助
自我中心的人类视频为模仿学习提供了可扩展的示范,但现有的语料库往往缺乏细粒度、时间局部化的动作描述或灵巧手部标注。我们提出了OpenEgo,一个多模态的自我中心操作数据集,具有标准化的手部姿态标注和意图对齐的动作原语。OpenEgo总计1107小时,涵盖六个公开数据集,覆盖600多个环境中的290个操作任务。我们统一了手部姿态布局,并提供了带时间戳的描述性动作原语。为验证其实用性,我们训练了语言条件模仿学习策略来预测灵巧手部轨迹。OpenEgo旨在降低从自我中心视频学习灵巧操作的障碍,并支持视觉-语言-动作学习中的可复现研究。所有资源和说明将在www.openegocentric.com发布。
英文摘要
Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.
发表机构
- Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系)
- Physical Automation, Inc.(Physical Automation 公司)
机构由 AI 辅助整理,请以论文原文为准。