arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Astronex-World 1.0:实时交互式世界模型基础

Astronex-World 1.0: Real-Time Interactive World Model Foundation

Xin Zhou, Cong Miao

arXiv 2609.20034首次发表:更新:

发表机构

Astronex Robotics; Nanjing University of Information Science and Technology(Astronex机器人公司; 南京信息工程大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Astronex-World 1.0提出可控视频世界模型,支持文本/图像到视频预测,采用双向与因果架构,五阶段训练,在WBench上超越更大模型,并支持具身智能后训练。

AI 中文摘要

我们提出了Astronex-World 1.0,一个开放的可控视频世界模型基础。给定文本提示(文本到视频)或初始观察(图像到视频),模型在帧对齐的相机轨迹、连续动作和具身标识符下预测未来视觉状态,并接受在展开(rollout)指定位置插入的文本事件。该系列提供了一个用于全上下文生成的双向模型和一个具有块因果注意力和跨块KV缓存的因果模型,用于持续生成,两者均基于Wan2.2-TI2V-5B先验构建。PRoPE注入相机内参和外参,而64维动作流调制每个Transformer层。五阶段训练路径开发了双向相机和动作控制,将骨干转换为块因果生成,蒸馏出少步学生模型,恢复混合域动态,并应用非对称DMD/DMD2分布匹配。因果模型以24 fps生成832x480视频。所有五个训练阶段均在两块NVIDIA L20 48 GB GPU上运行,因果模型在一块GPU上实时流式生成。它在WBench Navi上得分73.5,在WBench Full上得分70.0。在Full上,这个5B模型优于13.6B的LongCat-Video和14B的Helios,与22B的LTX-2.3相差不到1分,并优于YUME 1.5(后者从相同的5B先验在NVIDIA A100 GPU上进行后训练)。保留的动作输入和输出接口允许为具身智能和自动驾驶进行后训练。

英文摘要

We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.

CommentsTechnical report. 25 pages, 13 figures, 10 tables. Project page: https://world.astronex.com.cn ; Code: https://github.com/Astronex-Robotics/Astronex-World ; Weights: https://huggingface.co/Astronex-Lab/Astronex-World

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑