Zing-0.5:迈向具有实时联合动作与文本控制的可玩世界
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
浏览论文内容
中文总结 AI 辅助
Zing-0.5是一个5B参数的自回归世界模型,通过统一动作与文本条件化、事件级监督蒸馏和低成本流式推理,实现实时联合控制的可玩世界生成,在WBench导航上取得81.0分。
中文摘要 AI 辅助
我们推出了Zing-0.5,一个5B参数的自回归世界模型,专为可玩性设计:用户可以通过联合键盘和在线文本控制,探索生成的世界、影响正在展开的事件,并对由此产生的反馈做出响应。我们的方法汇集了三项技术贡献:(1)统一的动作与文本条件化,将幅度感知的键盘输入与时间对齐的文本指令以及联合标注的视频相结合,以在同一序列中学习导航和事件控制;(2)用于增量生成的事件级监督,利用在连接的多提示视频上训练的片段级教师模型,通过分布匹配蒸馏来监督块级因果学生模型;(3)低成本的实时交互,将四步生成与上下文保持的流式处理相结合,支持832×480分辨率下24 FPS的推理,估计服务器租赁成本约为每流分钟0.009美元。Zing-0.5在158个WBench导航案例中取得了81.0的总体得分和88.5的一致性得分。一个联合控制演示展示了在继续导航过程中无需重新启动生成即可实现文本引导的事件变更。我们发布了模型权重、推理代码以及Zing-SGLang服务实现,以支持可玩生成世界的进一步研究。
英文摘要
We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online text control. Our approach brings together three technical contributions: (1) Unified action and text conditioning, combining magnitude-aware keyboard inputs with temporally aligned text instructions and jointly annotated videos to learn navigation and event control within the same sequence; (2) Event-scale supervision for incremental generation, using a segment-level teacher trained on connected multi-prompt videos to supervise a block-level causal student through distribution-matching distillation; and (3) Low-cost real-time interaction, combining four-step generation with context-preserving streaming to support 832 x 480 inference at 24 FPS at an estimated server rental cost of approximately USD 0.009 per stream-minute. Zing-0.5 achieves an overall score of 81.0 and a consistency score of 88.5 across 158 WBench Navigation cases. A joint-control demonstration shows a text-directed event change during continued navigation without restarting generation. We release the model weights, inference code, and Zing-SGLang serving implementation to support further work on playable generated worlds.