arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Valerant:通过动作条件世界模型探索的自动可导航游戏地图生成器

Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration

Yiran Qiao, Feng Wang, Jing Ma

arXiv 2609.09418首次发表:更新:

发表机构

Case Western Reserve University; Johns Hopkins University(凯斯西储大学; 约翰斯·霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Valerant提出无需训练框架,将预训练动作条件世界模型与SLAM重建结合,从单张图像自动生成可导航3D游戏地图,减少人工创建成本。

AI 中文摘要

世界动作模型(WAMs)将预测性世界建模与动作生成相结合,使得预期的未来状态能够指导智能体行为。尽管WAMs正在快速推动具身人工智能的发展,但在游戏中,通用对应的模型在很大程度上仍未得到探索。现有的面向游戏的方法通常将动作条件世界模型与外部策略和奖励函数相结合,以实现类似WAM的决策,但这些方法主要运行在二维视觉观察空间中,并未实例化持久的3D几何结构。将这一范式扩展到3D游戏引入了独特的挑战。在自动驾驶和机器人技术中,物理环境独立于模型存在,提供了一个持久的3D世界,在其中可以执行选定的动作。游戏没有这样的外部基础;虚拟世界本身必须被实例化。大多数可玩游戏需要持久且可导航的空间,而3D游戏还额外需要支持移动和交互的显式几何结构。动作条件视频推演提供视觉观察,但不提供这种空间表示。我们提出了Valerant,一个无需训练的框架,将预训练的动作条件世界模型转变为WAM,用于探索和构建3D游戏地图。通过将预测性视觉推演与基于SLAM的空间重建和探索驱动的动作选择相结合,Valerant逐步将单张图像转变为持久的3D游戏地图。该框架将基于WAM的交互扩展到2D视觉模拟之外,并提供了一种减少3D游戏地图创建中人工努力的新方法。

英文摘要

World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in which selected actions can be executed. Games have no such external substrate; the virtual world itself must be instantiated. Most playable games require a persistent and navigable space, while 3D games additionally require explicit geometry that supports movement and interaction. Action-conditioned video rollouts provide visual observations but not this spatial representation. We present \textsc{Valerant}, a training-free framework that transforms a pretrained action-conditioned world model into a WAM for exploring and constructing 3D game maps. By coupling predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection, \textsc{Valerant} progressively transforms a single image into a persistent 3D game map. This framework extends WAM-based interaction beyond 2D visual simulation and offers a new approach to reducing manual effort in 3D game-map creation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑