arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.08269cs.CV

EgoX:从单个外视视频生成自视视频

EgoX: Egocentric Video Generation from a Single Exocentric Video

Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park, Junha Hyung, Jaegul Choo

首次发表 更新
浏览论文内容

中文总结 AI 辅助

EgoX通过轻量级LoRA适应和统一条件策略,从单个外视视频生成自视视频,实现几何一致且逼真的视频生成,具有高可扩展性和鲁棒性。

中文摘要 AI 辅助

自视感知使人类能够直接从自身视角体验和理解世界。将外视(第三人称)视频转换为自视(第一人称)视频开辟了沉浸式理解的新可能,但因极端的摄像机姿态变化和最小的视图重叠而极具挑战性。该任务要求忠实保留可见内容,同时以几何一致的方式合成未见区域。为此,我们提出了EgoX,一种从单个外视输入生成自视视频的新框架。EgoX利用大规模视频扩散模型的预训练时空知识,通过轻量级LoRA适应进行调整,并引入一种统一的条件策略,通过宽度和通道级拼接结合外视和自视先验。此外,一种几何引导的自注意力机制会优先关注空间相关的区域,确保几何一致性和高视觉保真度。我们的方法实现了连贯且逼真的自视视频生成,同时在未见过和真实世界视频上展示了强大的可扩展性和鲁棒性。

英文摘要

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibilities for immersive understanding but remains highly challenging due to extreme camera pose variations and minimal view overlap. This task requires faithfully preserving visible content while synthesizing unseen regions in a geometrically consistent manner. To achieve this, we present EgoX, a novel framework for generating egocentric videos from a single exocentric input. EgoX leverages the pretrained spatio temporal knowledge of large-scale video diffusion models through lightweight LoRA adaptation and introduces a unified conditioning strategy that combines exocentric and egocentric priors via width and channel wise concatenation. Additionally, a geometry-guided self-attention mechanism selectively attends to spatially relevant regions, ensuring geometric coherence and high visual fidelity. Our approach achieves coherent and realistic egocentric video generation while demonstrating strong scalability and robustness across unseen and in-the-wild videos.

发表机构

  • KAIST AI(韩国科学技术院AI)
  • Seoul National University(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑