arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Matrix-Game 3.5:利用补丁内存增强实时流式交互式世界模型

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

Runjia Qian, Zile Wang, Jihai Zhang, Kai Zou, Wei Yu, Jiaxing Li, Zexiang Liu, Yaokun Li, Fei Kang, Kaichen Huang, Mengyin An, Haobo Zhang, Biao Jiang, Jiahua Wang, Haofeng Sun, Yang Liu, Yangguang Li

arXiv 2608.29910首次发表:更新:

发表机构

Riemann Dynamics(黎曼动力学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Matrix-Game 3.5通过三项关键改进优化交互式世界模型,实现稳定长时序实时交互式生成,在多类任务上表现出色。

AI 中文摘要

交互式世界模型将视频生成从离线片段合成扩展到交互式虚拟世界的持续模拟,可应用于游戏、机器人、具身智能体和扩展现实(XR)领域。然而,实现稳定的长时序交互式生成仍具挑战,因为模型需同时保留场景几何、动态一致性和相机控制,同时支持实时自回归生成。在Matrix-Game 3.0的基础上,本文提出Matrix-Game 3.5,如图1所示,通过三项关键改进推进实时交互式世界生成,实现几何感知和长时序一致的模拟。其一,提出统一的几何感知内存框架,其补丁内存(patch-memory)和 tiled-PRoPE 组件不引入额外可学习参数,结合显式3D补丁检索与投影相机条件,实现几何一致的相机控制和准确的长时序场景回忆。其二,引入静态-动态解耦的世界表示,分别建模静态场景几何与动态主体,在长时序生成中保留几何一致性和主体身份。其三,开发两阶段渐进式实时蒸馏框架,通过感知流匹配(Perceptual Flow Matching)和基于课程的自展开动态模态分解(Self-Rollout DMD),将双向扩散模型转换为少步因果生成器,支持分钟级的实时交互式生成。大量实验表明,Matrix-Game 3.5在涵盖Unreal模拟环境、开放世界游戏和互联网视频的统一训练语料上,在长时序场景回忆、精确相机控制、主体一致性、提示驱动的世界生成以及稳定的实时开放世界交互方面表现出色。

英文摘要

Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.

Commentshttps://matrix-game-v3-5.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑