arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01127cs.CV

MiniWorld:从零开始普及视频世界模型的训练

MiniWorld: Democratizing the Training of Video World Models from Scratch

Yian Zhao, Ruochong Zheng, Hongcan Guo, Yu Yan, Jian Zhang, Jie Chen

首次发表
浏览论文内容

中文总结 AI 辅助

MiniWorld是从零开始训练流式视频世界模型的可复现框架,采用特定模型与训练策略,可在单台8-GPU服务器数天内完成训练,旨在降低训练门槛以推动相关研究。

中文摘要 AI 辅助

视频世界模型会根据历史观测值和控制信号预测未来观测值,通过自回归状态转换实现长时序生成。与主要捕捉视觉外观和运动的传统视频生成模型不同,视频世界模型学习智能体动作下环境演化的底层动态,为具身AI和交互式模拟提供基础。近期进展大多依赖通过后训练或蒸馏适配预训练视频生成模型,这些方法虽有效,但往往需要复杂的训练流程、大量计算资源,且存在双向预训练与因果流式推理的不匹配问题。近期研究表明,从零开始训练自回归视频世界模型是可行且可扩展的,但学界仍缺乏一种轻量、透明、完全可复现、可在适度计算资源下端到端训练的基线。我们提出MiniWorld,这是一个用于从零开始训练流式视频世界模型的可复现框架。MiniWorld采用块因果视频扩散Transformer,在预训练视频VAE的潜在空间中通过流匹配进行训练;它基于Diffusion Forcing,采用分块非递减噪声调度和两阶段持续训练,以提升时序建模能力和稳定性。推理时,MiniWorld结合滚动KV缓存与流水线异步去噪,在有限计算下实现高效流式生成。整个模型可在单台8-GPU服务器上数天内完成训练,我们发布训练和推理代码库及预训练检查点,希望MiniWorld能推动视频世界建模的未来研究。

英文摘要

Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions. Unlike conventional video generation models that primarily capture visual appearance and motion, video world models learn the underlying dynamics governing environment evolution under agent actions, providing a foundation for embodied AI and interactive simulation. Recent progress has largely relied on adapting pretrained video generation models through post-training or distillation. Although effective, these approaches often require complex training pipelines, substantial computational resources, and suffer from the mismatch between bidirectional pretraining and causal streaming inference. Recent studies have shown that training autoregressive video world models from scratch is feasible and scalable. However, the community still lacks a lightweight, transparent, and fully reproducible baseline trainable end-to-end with modest computational resources. We present MiniWorld, a reproducible framework for training streaming video world models from scratch. MiniWorld employs a block-causal Video Diffusion Transformer trained with Flow Matching in the latent space of a pretrained Video VAE. Building on Diffusion Forcing, it adopts a chunk-wise non-decreasing noise schedule and two-stage continued training to improve temporal modeling and stability. During inference, MiniWorld combines a rolling KV cache with pipelined asynchronous denoising for efficient streaming generation under bounded computation. The entire model can be trained within several days on a single 8-GPU server. By releasing the training and inference codebase and pretrained checkpoints, we hope MiniWorld will facilitate future research on video world modeling.

发表机构

  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑