AlayaVista:从全景状态到透视视频的流式世界建模
AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video
浏览论文内容
中文总结 AI 辅助
AlayaVista提出解耦全景演化与透视合成的流式世界模型,利用全景先验和潜在渲染实现相机可控视频生成,并构建MUGEN数据集支撑训练。
中文摘要 AI 辅助
交互式视频世界模型必须在相机运动下保持广泛的场景上下文,同时以低延迟生成高保真观测。现有方法面临表示上的权衡:透视模型在局部视图上操作,并必须在长时间展开中保留屏幕外内容,而更广泛的空间覆盖通常通过合成全球视频或构建显式3D表示来获得。受全局上下文与选择性局部敏锐度在视觉感知中互补作用的启发,我们提出了AlayaVista,一种相机可控的流式视频世界模型,它将全景世界演化与透视观测合成解耦。给定单张透视图像,AlayaVista使用预训练的全景扩展模型构建360度场景先验,然后将场景演化为相机条件下的全景潜在状态。一个潜在视口渲染器将该状态映射到所请求的透视视频潜在表示,而透视细化器则恢复细节、抑制伪影并执行超分辨率。为了支持高效流式处理,我们将全景生成器适配为分块自回归生成,并将全景生成和透视细化均蒸馏为少步过程。为了提供该设计所需的监督,我们构建了MUGEN,一个大规模真实世界全景视频数据集,包含1,318小时分辨率至少为4K的视频,并附有丰富的语义和几何标注。
英文摘要
Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selective local acuity in visual perception, we present AlayaVista, a camera-controllable streaming video world model that decouples panoramic world evolution from perspective observation synthesis. Given a single perspective image, AlayaVista constructs a 360-degree scene prior using a pretrained panorama expansion model and then evolves the scene as a camera-conditioned panoramic latent state. A latent viewport renderer maps this state to the requested perspective video latents, while a perspective refiner restores details, suppresses artifacts, and performs super-resolution. To support efficient streaming, we adapt the panoramic generator to chunk-autoregressive generation and distill both panoramic generation and perspective refinement into few-step processes. To provide the supervision required by this design, we construct MUGEN, a large-scale real-world panoramic video dataset containing 1,318 hours of videos at resolutions of at least 4K, together with rich semantic and geometric annotations.
发表机构
- Alaya Lab(阿拉亚实验室)
- Beijing Institute of Technology(北京理工大学)
- The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。