WorldWander:在视频生成中连接自我中心和外中心世界
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
浏览论文内容
中文总结 AI 辅助
研究聚焦视频生成中自我中心与外中心世界转换,提出WorldWander框架,基于视频扩散变换器,整合上下文视角对齐与协作位置编码,精心策划数据集,实验证明该框架在视角同步、角色一致性和泛化能力上表现卓越,设定了新基准。
中文摘要 AI 辅助
视频世界模型的最新进展实现了具有自由导航的交互式环境,使得第一人称(自我中心)和第三人称(外中心)视角之间的转换变得越发重要。然而,现有研究集中于单向的从外中心到自我中心的转换,忽视了参考引导的外中心视角合成。这种能力对游戏和具身人工智能应用至关重要。为此,我们提出了WorldWander,一个为视频生成中自我中心和外中心世界之间的转换量身定制的上下文学习框架。基于先进的视频扩散变换器,WorldWander整合了上下文视角对齐和协作位置编码,以对跨视角同步和角色一致性进行建模。为支持我们的任务,我们精心策划了EgoExo-8K,这是一个包含来自合成和现实世界场景的同步自我中心-外中心三元组的动态且场景丰富的数据集。实验表明,WorldWander实现了卓越的视角同步、角色一致性和泛化能力,为自我中心-外中心视频转换设定了新的基准。
英文摘要
Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However, existing studies focus on unidirectional exocentric-to-egocentric translation, overlooking reference-guided exocentric perspective synthesis. This capability is crucial for gaming and embodied AI applications. Motivated by this, we present WorldWander, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. Building upon advanced video diffusion transformers, WorldWander integrates (i) In-Context Perspective Alignment and (ii) Collaborative Position Encoding to model cross-view synchronization and character consistency. To support our task, we curate EgoExo-8K, a dynamic and scene-rich dataset containing synchronized egocentric-exocentric triplets from both synthetic and real-world scenarios. Experiments demonstrate that WorldWander achieves superior perspective synchronization, character consistency, and generalization, setting a new benchmark for egocentric-exocentric video translation.
发表机构
- Show Lab, National University of Singapore(新加坡国立大学Show实验室)
机构由 AI 辅助整理,请以论文原文为准。