换位思考:从外中心视角视频中提取自我中心视角
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos
浏览论文内容
中文总结 AI 辅助
本研究针对外中心到自我中心的跨视图视频转换问题,提出两阶段生成框架Exo2Ego,构建了含多数据集的基准,实验证明其在合成质量与泛化性上优于基线。
中文摘要 AI 辅助
我们研究外中心视角到自我中心视角的跨视图转换任务,该任务旨在基于从第三人称(外中心)视角拍摄动作执行者的视频记录,生成第一人称(自我中心)视角的画面。为此,我们提出了名为Exo2Ego的生成框架,将转换过程解耦为两个阶段:高层结构转换,显式促进外中心视角与自我中心视角之间的跨视图对应关系;以及基于扩散模型的像素级生成,引入手部布局先验以提升生成的自我中心视角画面的保真度。为推动该领域的未来发展,我们构建了一个全面的外到内跨视图转换基准,包含来自H2O、Aria Pilot和Assembly101三个公开数据集的多样化同步自我-外中心桌面活动视频对。实验结果验证,Exo2Ego能够生成具有清晰手部操作细节的照片级真实视频效果,在合成质量和对新动作的泛化能力方面均优于多个基线模型。
英文摘要
We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this end, we propose a generative framework called Exo2Ego that decouples the translation process into two stages: high-level structure transformation, which explicitly encourages cross-view correspondence between exocentric and egocentric views, and a diffusion-based pixel-level hallucination, which incorporates a hand layout prior to enhance the fidelity of the generated egocentric view. To pave the way for future advancements in this field, we curate a comprehensive exo-to-ego cross-view translation benchmark. It consists of a diverse collection of synchronized ego-exo tabletop activity video pairs sourced from three public datasets: H2O, Aria Pilot, and Assembly101. The experimental results validate that Exo2Ego delivers photorealistic video results with clear hand manipulation details and outperforms several baselines in terms of both synthesis quality and generalization ability to new actions.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- FAIR at Meta(Meta FAIR实验室)
机构由 AI 辅助整理,请以论文原文为准。