arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReViV:从单目自我中心视频中重建4D中的观看者和视图

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang

arXiv 2607.17790首次发表:更新:

发表机构

ETH Zurich; Delft University of Technology; Microsoft(苏黎世联邦理工学院; 代尔夫特理工大学; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在从单目自我中心视频重建4D中的观看者和视图,提出ReViV框架,将任务建模为学习多模态信号联合概率分布,由掩码生成自我中心变压器驱动,在多基准测试中展现出高精度和效率,还保持了竞争力强的自我中心深度估计,且代码模型开源。

AI 中文摘要

自我中心设备,如可穿戴前置摄像头,为捕捉人类观看者与周围环境之间的持续交互提供了独特视角。因此,非常需要一个能够重建这种4D表示的整体高效多模态模型。然而,现有方法往往依赖辅助输入,将场景感知和人类自我运动建模视为相互独立的问题,且推理时间长。为解决这些局限,我们提出ReViV,首个从单目RGB视频中提取观看者和视图动态的整体自我中心4D重建统一框架。我们将任务表述为学习多模态信号的全联合概率分布,由掩码生成自我中心变压器驱动,在单一前馈架构中运行,以快速推理速度同时重建观看者和视图的时间一致4D重建。在多个基准测试上的大量实验表明,ReViV在整体自我身体、手部和注视重建、相机跟踪方面达到了当前最优的精度和效率,在不依赖繁重特定任务先验的情况下保持了极具竞争力的自我中心深度估计。代码和模型已完全开源。

英文摘要

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.

CommentsAccepted to ECCV 2026. The first two authors contributed equally, and their author order is interchangeable

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑