arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GameWAM:面向电子游戏的世界动作模型

GameWAM: A World Action Model for Video Games

Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li

arXiv 2608.26200首次发表:更新:

发表机构

Fudan University; LIGHTSPEED; Tsinghua Shenzhen International Graduate School(复旦大学; 光速(企业名); 清华大学深圳国际研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GameWAM是首个用于原生闭环游戏玩法和GUI控制的世界动作模型,通过并行生成过程实现世界-动作联合学习,实验表明其以更少原生动作达到竞争力任务成功率,还发现了LASI失效模式。

AI 中文摘要

现代电子游戏结合了第一人称感知、快速视觉变化、持久的世界状态以及异构的原生控制方式。现有游戏智能体将视觉和任务上下文直接映射到动作,但缺乏显式的世界动态建模;而交互式游戏世界模型会根据给定动作预测视觉未来,但无法作为任务策略。世界动作模型(WAM)将这些目标统一起来,但在电子游戏的动态性和开放式交互场景下仍未得到充分探索。我们提出了GameWAM,据我们所知,这是首个用于原生闭环游戏玩法和图形用户界面(GUI)控制的WAM。GameWAM通过具有块因果条件和流匹配的并行视觉与动作生成过程,联合生成未来的视觉观测结果和可执行的键盘-鼠标轨迹。为支持世界-动作联合学习,我们构建了同步的游戏玩法和GUI轨迹。为处理异构的原生控制,GameWAM在每个动作步骤预测游戏玩法/GUI模式,并使用模式特定的预测分布和连续动作归一化来生成动作。对于长时序交互,块周期控制会在承诺的时域之外进行预测,仅执行短的动作前缀,并从新观测中重新规划,同时细粒度的周期内上下文和分层的跨周期历史保持时间连续性。实验表明,与对比智能体相比,GameWAM以更少的已执行原生动作实现了具有竞争力的任务成功率。我们进一步发现了低频动作源印记(LASI)现象:在固定条件下,采样动作源的低频分量会系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。项目页面可通过此https URL访问。

英文摘要

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World-Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open-ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed-loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching. To support joint world-action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode per action step and generates actions with mode-specific prediction distributions and continuous-action normalization. For long-horizon interaction, block-cycle control coordinates prediction, execution, and temporal context: it predicts beyond the committed horizon, executes short action blocks, replans from new observations, and hierarchically structures context from fine-grained within-cycle history to persistent cross-cycle history. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.

Comments44 pages, 23 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑