arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Grounded-Exo2Ego:用于鲁棒外部视角到自我视角视频生成的结构化语义 grounding

Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Shengze Wang, Michael Stengel, Tianye Li, Seonwook Park, Amrita Mazumdar, Koki Nagano, Alex Trevithick, Shalini De Mello

arXiv 2608.20534首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Grounded-Exo2Ego框架,通过双分支视频扩散模型、相机重定位算法和自动合成数据引擎,在EgoExo4D数据集上大幅提升了外部视角到自我视角视频生成的性能。

AI 中文摘要

从单段外部视角(exocentric)视频生成自我视角(egocentric)视频是针对增强现实(AR)、虚拟现实(VR)和物理人工智能的新兴重要研究方向。与传统的新视角合成相比,外部视角到自我视角的生成任务难度显著更高,因为在极端视角变化和存在大量不可观测区域的情况下,标准几何条件会变得极不可靠。本文提出Grounded-Exo2Ego,这是一种在架构和数据层面解决上述挑战的原则性框架。在架构层面,Grounded-Exo2Ego是一种双分支视频扩散模型,它将几何锚定分支(该分支以3D重建的渲染结果作为生成条件)与一种新型语义 grounding 分支相结合;该分支超越了主流的基于几何的方法,通过基于物体级上下文合成具有挑战性的区域来提升生成质量。此外,研究发现被忽视的相机-重建错位问题会严重损害外部视角到自我视角的学习,因此本文引入一种相机重定位算法来解决该问题,并在所有指标上大幅提升生成质量。本文还开发了一种全自动合成数据引擎,可在程序生成的环境中生成并渲染绑定骨骼的3D角色。在具有挑战性的EgoExo4D数据集上的评估表明,本文方法在所有指标上均大幅优于近期的最先进方法,详细的 ablation 实验验证了架构和数据层面各项贡献的改进效果。

英文摘要

Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego generation is a significantly harder task because the standard geometric conditioning becomes highly unreliable under extreme view changes and large unobservable regions. We present Grounded-Exo2Ego, a principled framework that addresses these challenges at both the architectural and data levels. Architecturally, Grounded-Exo2Ego is a dual-branch video diffusion model that couples a geometric anchoring branch, which conditions the generation on the rendering of a 3D reconstruction, with a novel semantic grounding branch, which goes beyond the prevailing geometry-based approach and improves quality by synthesizing challenging regions based on object-level context. Additionally, we found that the overlooked issue of camera-reconstruction misalignment severely undermines exo-to-ego learning. We thus introduce a camera re-localization algorithm that resolves this issue and substantially improves quality across all metrics. We further develop a fully automated synthetic data engine that generates and renders rigged 3D characters in procedurally generated environments. Evaluation on the challenging EgoExo4D dataset shows that our method outperforms recent state-of-the-art approaches by large margins across all metrics. Detailed ablations validate improvements from each of our contributions at both the data and architectural level.

Commentswebsite url: https://research.nvidia.com/labs/amri/projects/grounded-exo2ego/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑