arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.02886cs.CV

SolarWM:面向长时序视频世界模型的开放数据与可扩展训练

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu H… 展开作者

Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

SolarWM是面向长时序视频世界模型的开放基础框架,通过多源数据引擎与适配框架实现异构数据训练,实例化多款5B至33B模型,仅用5秒序列训练即可支持长时序实时交互,为相关研究提供可复扩展的基础。

中文摘要 AI 辅助

我们推出SolarWM,这是一个完全开放的基础框架,用于构建从数据准备到长时序推理的交互式视频世界模型。跨异构数据源和视频主干进行训练颇具挑战:不同数据集在时间尺度、相机几何结构、视觉质量、运动及字幕风格上存在差异,而视频生成器则采用不同的表示方式和架构。因此,直接混合数据和采用特定于模型的实现会产生不一致的监督信号,且使得结果难以复现和比较。SolarWM通过可重构的多源数据引擎和原生主干适配框架解决了这种耦合问题。该引擎将来自10个数据集的143万个规范片段转换为统一的、帧对齐的契约,涵盖视觉观测、度量相机几何结构、字幕、质量元数据、选择决策及来源,同时将源处理与混合构建解耦。在共享的相机条件、训练和推理接口下,我们基于Wan2.2、LTX-2.5和MiniMax-H3实例化了四个5B至33B的模型,同时保留了它们的原生表示和目标。统一的三阶段方案结合了双向适配、教师强制自回归初始化以及分布匹配蒸馏。所得的因果模型在仅用5秒序列训练后,即可在从几分钟到几小时的 rollouts 上实现实时交互。通过发布生成的数据、流水线、方案、权重和框架,SolarWM为交互式世界模型研究提供了可复现且可扩展的基础。

英文摘要

We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.

发表机构

  • CUHK-SZ(香港中文大学(深圳))
  • SLAI(深圳先进人工智能研究院)
  • NUS(新加坡国立大学)
  • CUHK(香港中文大学)
  • HKUST(香港科技大学)
  • HKUST-GZ(香港科技大学(广州))
  • NVIDIA(英伟达公司)
  • UCLA(加利福尼亚大学洛杉矶分校)
  • MSRA(微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑