arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ACE-Data-0:以人为中心的环境捕获作为具身数据引擎

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

arXiv 2607.28625首次发表:更新:

发表机构

S-Lab, Nanyang Technological University; ACE Robotics(新加坡南洋理工大学S-Lab; ACE机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以人为中心的具身数据引擎ACE,构建含150小时数据的ACE-Data-0数据集,引入分层基准,暴露现有方法缺陷,为具身AI等领域提供基础。

AI 中文摘要

具身智能面临着根本性的数据瓶颈:模型需要捕捉人类在追求目标时,第一人称感知、全身运动、灵巧操作、物体状态、声音和触觉如何随时间协同演化。现有数据集将这种体验按视角、模态或空间尺度拆分,导致完整的感知-行动循环仅被部分观测。我们提出Ambient Capture Engine(ACE,环境捕获引擎),这是一种以人为中心的数据引擎,可将真实家庭环境转换为空间校准、时间同步的录制工作室。ACE在两个互补尺度运行:桌面级配置解析手-物体操作,房间级配置捕捉全身运动、 locomotion( locomotion译为“移动”)以及带家具家庭内的交互。ACE记录自我中心和多视角外部中心视频、全身与关节手运动、物体几何与6自由度(6-DoF)轨迹、音频及触觉信号,形成统一多感官流。利用ACE,我们构建ACE-Data-0,包含200类任务、50名参与者在2个环境中完成的150小时时长、1700万视频帧,共75000个交互片段。该数据集涵盖原子级操作、长时程家庭活动链、人与场景交互,且通过目标级而非分步指令保留自然行为变异。我们还引入从信号到场景组件再到交互的分层基准,对当前最优方法的评估显示其在接触、遮挡、自我运动和长时程条件下存在显著差距。ACE-Data-0提供带对齐感知、运动学和接触监督的同步人类演示,为模仿学习、世界模型、视觉-语言-行动系统及具身AI提供可扩展基础。

英文摘要

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

CommentsProject Page: https://ace-data-engine.github.io/ACE-Data-0/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑