发表机构
Meta Reality Labs; ETH Zürich(元现实实验室; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对头戴式设备人体运动捕捉,提出EgoExoMoCap分布式框架,联合自我与外部中心多模态信号,利用头部等跟踪信号及DINOv3特征,能在复杂场景中稳健重建运动,突破了两种范式孤立的局限。
AI 中文摘要
头戴式设备的人体运动捕捉为获取现实世界的人体运动和交互数据提供了一种可扩展的方式,这对具身人工智能和VR/AR应用至关重要。现有方法要么专注于自我中心身体跟踪,要么专注于外部中心跟踪,这两种范式大多孤立探索。本文提出一种新颖的分布式框架,联合利用自我和外部中心多模态信号从头戴式设备进行人体运动估计。与传统系统不同,该方法简单如两人各戴一副智能眼镜。它利用头部(可能还有手腕)跟踪信号准确估计3D世界中的全局运动,并结合基于DINOv3的上下文感知图像特征以在有噪声和遮挡的情况下实现鲁棒性。在两个野外数据集上的大量实验表明,该方法即使在具有挑战性的场景中也能稳健地重建运动。
英文摘要
Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer's surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.
CommentsAccepted by ECCV 2026, Project page and code: https://siplab.org/projects/EgoExoMoCap