arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2511.20186cs.CV

Exo2EgoSyn: 解锁用于外视图到内视图视频合成的基础视频生成模型

Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

Mohammad Mahdi, Yuqian Fu, Nedko Savov, Jiancheng Pan, Danda Pani Paudel, Luc Van Gool

首次发表
浏览论文内容

中文总结 AI 辅助

Exo2EgoSyn通过三个模块实现从第三人称视角生成高保真内视图视频,提升跨视角视频合成能力。

中文摘要 AI 辅助

基础视频生成模型如WAN 2.2表现出强大的文本和图像条件合成能力,但仍然受限于同一视角生成设置。在本文中,我们介绍了Exo2EgoSyn,这是WAN 2.2的改进版本,解锁了外视图到内视图(Exo2Ego)跨视角视频合成。我们的框架由三个关键模块组成。Ego-Exo视图对齐(EgoExo-Align)在外视图和内视图的首帧表示之间强制潜在空间对齐,将生成空间从给定的外视图转向内视图。多视角外视图视频条件(MultiExoCon)将多视角外视图视频聚合为统一的条件信号,将WAN2.2扩展到其原始单图像或文本条件之外。此外,姿态感知的潜在注入(PoseInj)将相对外视图到内视图相机姿态信息注入到潜在状态中,引导在不同视角下的几何感知合成。这些模块共同实现了从第三人称观察生成高保真的内视图视频,而无需从头开始重新训练。在ExoEgo4D上的实验表明,Exo2EgoSyn显著提高了Ego2Exo合成,为使用基础模型实现可扩展的跨视角视频生成铺平了道路。源代码和模型将公开发布。

英文摘要

Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adaptation of WAN 2.2 that unlocks Exocentric-to-Egocentric(Exo2Ego) cross-view video synthesis. Our framework consists of three key modules. Ego-Exo View Alignment(EgoExo-Align) enforces latent-space alignment between exocentric and egocentric first-frame representations, reorienting the generative space from the given exo view toward the ego view. Multi-view Exocentric Video Conditioning (MultiExoCon) aggregates multi-view exocentric videos into a unified conditioning signal, extending WAN2.2 beyond its vanilla single-image or text conditioning. Furthermore, Pose-Aware Latent Injection (PoseInj) injects relative exo-to-ego camera pose information into the latent state, guiding geometry-aware synthesis across viewpoints. Together, these modules enable high-fidelity ego view video generation from third-person observations without retraining from scratch. Experiments on ExoEgo4D validate that Exo2EgoSyn significantly improves Ego2Exo synthesis, paving the way for scalable cross-view video generation with foundation models. Source code and models will be released publicly.

↑