arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界中的世界:用世界模型探索世界

World in World: Explore the World with World Models

Chenxi Song, Yanming Yang, Chi Zhang

arXiv 2609.11548首次发表:更新:

发表机构

Westlake AGI Lab(西湖大学通用人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出World in World,一种无需训练、仅推理时使用的接口,通过将异构控制证据转化为视觉状态并利用冻结视频模型的注意力机制,实现相机控制重渲染、长时程重访和运动迁移。

AI 中文摘要

自回归视频世界模型能够实现交互式、长时程的探索,但灵活的控制仍然具有挑战性。从新的视角探索源视频,要求生成的展开(rollout)与记录的事件保持同步,将观察到的内容放置在请求的视角中,合理地补全新暴露的区域,并在重新访问时恢复先前生成的外观。现有方法通常通过特定任务的模块或额外的训练来解决这些需求。我们提出了“世界中的世界”(World in World),一种无需训练、仅推理时使用的接口,它将异构的控制证据转换为带有相机和时间标签的干净视觉状态,并通过冻结的因果视频模型的原生自注意力机制进行读取。这些证据包括源视频观测、目标视角场景投影、指导新暴露主体区域补全的几何渲染,以及超出滚动缓存(rolling cache)的检索到的生成状态。每个证据源都带有令牌级别的支持及其自身的可用性调度。一个对应关系路由器(correspondence router)将持久的点身份与几何相结合,以建立令牌对应关系,引导受支持的查询匹配源视频令牌。随后,证据级注意力CFG(EWA)利用同一去噪前向传递中的注意力响应,独立调节每个辅助通道的额外贡献。该共享接口支持相机控制的重新渲染、长时程重访以及人体运动迁移,且使用相同的冻结骨干网络。我们在多种视角变化下的相机控制视频重新渲染任务上评估了“世界中的世界”,评估了感知质量、时间一致性和相机跟随准确性。

英文摘要

Autoregressive video world models enable interactive, long-horizon exploration, but flexible control remains challenging. Exploring a source video from new viewpoints requires the generated rollout to remain synchronised with the recorded event, place observed content in the requested view, plausibly complete newly exposed regions, and recover previously generated appearance on revisits. Existing methods typically address these requirements through task-specific modules or additional training. We present World in World, a training-free inference-time interface that converts heterogeneous control evidence into camera- and time-labelled clean visual states, which are read through the native self attention of a frozen causal video model. The evidence comprises source-video observations, target-view scene projections, geometry renderings that guide completion of newly exposed subject regions, and retrieved generated states beyond the rolling cache. Each evidence source carries token-level support and its own availability schedule. A correspondence router combines persistent point identities with geometry to establish token correspondences, guiding supported queries towards matching source-video tokens. Evidence-wise attention CFG (EWA) then independently regulates each auxiliary channel's additional contribution using attention responses from the same denoising forward pass. The shared interface supports camera-controlled rerendering, long-horizon revisiting, and human-motion transfer with the same frozen backbone. We evaluate World in World on camera-controlled video rerendering under diverse viewpoint changes, assessing perceptual quality, temporal consistency, and camera-following accuracy.

CommentsProject Page: https://chenxi-song.github.io/worldinworld

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑