发表机构
CUHK; Tencent PCG; FDU; Shanghai AI Laboratory; HKUST(香港中文大学; 腾讯平台与内容事业群; 复旦大学; 上海人工智能实验室; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ForgeWM 是一种渐进式框架,通过多阶段技术将双向动作条件视频生成器转化为少步骤世界模型,在 Minecraft 和 FPS 游戏任务中实现了更优的可控少步骤视频生成性能。
AI 中文摘要
动作条件视频世界模型需要低延迟的因果生成以及对游戏原生控制的可靠响应。尽管因果蒸馏能够实现单步或少步骤视频合成,但将其扩展到交互式世界模型仍具挑战性,因为离散键盘状态与连续鼠标运动必须在因果训练和自回归展开期间,与时间压缩的潜在块保持对齐。我们提出 ForgeWM,这是一个渐进式框架,通过领域适应、教师强制因果训练、因果一致性蒸馏以及与双向教师的在线策略分布匹配,将双向动作条件视频生成器转化为高效的少步骤世界模型。由此产生的预算专用学生模型在 1、2 和 4 步的稳态去噪预算下运行。ForgeWM 还支持双路径部署协议,将延迟关键型交互与可选的回放时细化相结合,其中单步学生模型会重新加噪并细化其保存的草稿。在配对的 Minecraft 轨迹上,ForgeWM 在成像质量、参考对齐运动轮廓一致性、动作符号准确率和鼠标控制准确率方面领先于所有评估系统,同时实现最低的参考 LPIPS;相同的四阶段方案可迁移到游戏手柄控制的 FPS 游戏玩法。回放时细化可达到四步参考质量,同时与经验轨迹的接近度约为从噪声中再生的三倍。这些结果证明了 ForgeWM 在可控少步骤视频生成方面的有效性。
英文摘要
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.