arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从第一视角视频感知世界与自身

Seeing the World and the Self from Egocentric Video

Kai Guan, Minchao Jiang, Ruichen WangLi, Wentao Zhu, Lei Zhang

arXiv 2609.01276首次发表:更新:

发表机构

The Hong Kong Polytechnic University; Eastern Institute of Technology(香港理工大学; 宁波工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对第一视角视频的三维感知难题,提出统一框架RESELF,结合几何重建与运动生成,在多项任务上优于现有最优方法,相关资源将公开。

AI 中文摘要

从第一视角视频实现完整的三维感知,需要在统一的度量坐标系中恢复周围场景及佩戴者的全身运动。现有方法通常分别处理场景重建与运动估计:场景重建方法忽略佩戴者,而运动估计方法缺乏明确的场景几何结构,且往往依赖外部轨迹。联合恢复极具挑战性,因为两项任务存在不对称的可见性,且需要不同的预测范式:高度可见的场景支持确定性几何回归,而严重被遮挡的人体则需要生成式运动推理。为此,我们提出RESELF(即“重建场景与自身”),这是一个统一框架,将确定性度量几何重建与几何条件运动生成相结合。RESELF通过逐帧尺度和相对位姿一致性目标,将在大规模第三视角数据上预训练的几何基础模型适配至第一视角视频。得到的相机轨迹与潜在几何特征会为扩散模型提供条件,以恢复佩戴者的运动。后续的闭环运动学反馈阶段会进一步优化相机头部,同时保留重建的场景几何结构。为支持训练与评估,我们从EgoExo4D中整理出EE4D-JSM数据集,该数据集对齐了第一视角视频、稀疏度量场景几何、相机轨迹及全身运动标注。实验表明,RESELF在深度估计、相机跟踪和全身运动估计任务上,均优于针对各单项任务设计的现有最优方法。代码、模型及数据集将在该httpsURL公开。

英文摘要

Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods lack explicit scene geometry and often depend on external trajectories. Joint recovery is challenging because the two tasks exhibit asymmetric visibility and require different prediction paradigms. The largely visible scene supports deterministic geometric regression, whereas the severely occluded body requires generative motion inference. We therefore propose RESELF (REconstructing the Scene and the sELF), a unified framework that couples deterministic metric geometry reconstruction with geometry-conditioned motion generation. RESELF adapts a geometry foundation model pre-trained on large-scale exocentric data to egocentric video using frame-wise scale and relative-pose consistency objectives. The resulting camera trajectory and latent geometric features condition a diffusion model that recovers the wearer's motion. A subsequent closed-loop kinematic feedback stage further refines the camera head while preserving the reconstructed scene geometry. To support training and evaluation, we curate EE4D-JSM from EgoExo4D by aligning egocentric video, sparse metric scene geometry, camera trajectories, and full-body motion annotations. Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation. Code, models, and datasets will be available at https://ka1guan.github.io/RESELF/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑