arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2509.05513cs.CVcs.AIcs.RO

OpenEgo:面向灵巧操作的大规模多模态自我中心数据集

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

Ahad Jawaid, Yu Xiang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

OpenEgo是一个大规模多模态自我中心操作数据集,含1107小时视频、290个任务,提供标准化手部姿态和动作原语,用于降低灵巧操作模仿学习门槛并支持可复现研究。

中文摘要 AI 辅助

自我中心的人类视频为模仿学习提供了可扩展的示范,但现有的语料库往往缺乏细粒度、时间局部化的动作描述或灵巧手部标注。我们提出了OpenEgo,一个多模态的自我中心操作数据集,具有标准化的手部姿态标注和意图对齐的动作原语。OpenEgo总计1107小时,涵盖六个公开数据集,覆盖600多个环境中的290个操作任务。我们统一了手部姿态布局,并提供了带时间戳的描述性动作原语。为验证其实用性,我们训练了语言条件模仿学习策略来预测灵巧手部轨迹。OpenEgo旨在降低从自我中心视频学习灵巧操作的障碍,并支持视觉-语言-动作学习中的可复现研究。所有资源和说明将在www.openegocentric.com发布。

英文摘要

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.

发表机构

  • Department of Computer Science, The University of Texas at Dallas(德克萨斯大学达拉斯分校计算机科学系)
  • Physical Automation, Inc.(Physical Automation 公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑