H3-World:将语言理解转化为世界控制
H3-World: Turning Language Understanding into World Control
浏览论文内容
中文总结 AI 辅助
H3-World是复用大型视频生成器语义表示的高效框架,仅需少量样本与参数,即可将语言接口转化为精确的交互式世界控制,且能泛化到未见场景。
中文摘要 AI 辅助
我们提出H3-World,这是一个高效框架,可将330亿参数的MiniMax-H3视频生成器转化为交互式世界模型。我们的核心发现是,随着大型视频生成器能力的提升,语言正逐渐成为一种自然的控制接口。例如,MiniMax-H3已支持通过自然语言指令对角色行为和相机运动进行零样本控制。在此基础上,H3-World将这种粗糙的语言接口转化为精确的、具有时间基础的世界控制,且无需引入专用动作模块。具体而言,我们将每个动作表示为角色指令与相机指令的结构化组合,并使其与对应时间的视频隐变量对齐。为使控制在时间上精确,我们进一步引入时间注意力路由,该机制将每条指令限制在其预期的时间区间内,减少动作间的控制泄漏。重要的是,H3-World直接复用了大规模视频预训练过程中学习到的语义表示,仅需轻量级适配。仅使用8000个游戏玩法样本、10000步LoRA优化,以及0.199%的可训练参数,H3-World即可实现有效的角色与相机控制,同时保持强大的生成质量,还能泛化到未见场景。这些结果表明,大型视频生成器中涌现的控制能力可被高效转化为交互式世界控制。
英文摘要
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.
发表机构
- Tencent(腾讯)
- National University of Singapore(新加坡国立大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。