所有人的全身姿态跟踪
Everybody Tracking Every Body
- University of California, Irvine(加利福尼亚大学欧文分校)
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对多互动个体的第一人称视角3D人体姿态估计问题,提出基于扩散的融合方法,结合第一人称头部运动与外参观测,在多人数据集上提升了姿态精度。
AI中文摘要:
我们研究的问题是:在存在集中式协调的情况下,从多个互动个体的第一人称视角估计他们的3D人体姿态。每个个体佩戴一台相机,用于记录第一人称视频和IMU数据。通过VIO SLAM处理该视频,可实现对每台第一人称相机在空间中的高质量跟踪。某一个体的第一人称视角会为其他人提供第三人称观测,不过这些外参观测是稀疏、间歇性的,且由于相机和主体都在移动,其可靠性波动极大。为整合这些同步数据流,我们提出一种基于扩散的方法,该方法融合了基于第一人称相机运动得出的头部运动姿态估计,以及外参姿态观测,融合过程以观测内容和可靠性为条件。我们的模型在单人动作捕捉数据与多人视频的混合数据集上进行训练,以学习人体运动轨迹和视频观测可靠性的丰富先验知识。在具有挑战性的多人数据集上开展的评估表明,与仅基于运动的基线方法、仅基于视觉的基线方法相比,我们的融合方法在绝对姿态精度和相对姿态精度方面均有所提升。
英文摘要:
We address the problem of 3D body pose estimation of multiple interacting people from their egocentric views with centralized coordination. Each individual wears a camera recording egocentric video and IMU data. Processing this video with VIO SLAM provides high-quality tracking of each egocentric camera through space. The first-person view from one individual provides third-person observations of other people, although these exocentric observations are sparse, intermittent, and of highly variable reliability as both cameras and subjects move. To integrate these synchronized data streams, we propose a diffusion-based approach that fuses estimates of pose based on head motion derived from egocentric camera motion with exocentric pose observations, conditioning on both observation content and reliability. Our model is trained on a mixture of single-person motion-capture data and multi-person video in order to learn rich priors for body motion trajectories and video observation reliability. Evaluation on challenging multi-person datasets suggests our fusion approach improves over motion-only and vision-only baselines in terms of both absolute and relative pose accuracy.