arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DashVMC:几何冲刺中的实时离散世界模型控制

DashVMC: Real-Time Discrete World Model Control in Geometry Dash

Florent Tariolle, Florian Yger

arXiv 2609.40003首次发表:更新:

AI 中文总结

DashVMC利用两小时游戏数据学习紧凑世界模型,通过BC初始化并用PPO在冻结回滚中优化,实现实时控制,在官方和社区关卡上超越基线,并维持60赫兹运行。

AI 中文摘要

世界模型智能体通常在可以等待策略的模拟器中进行评估;实时游戏则施加了相反的约束,要求在下一帧之前完成捕获、预测和动作。我们提出了DashVMC,它从大约两小时的录制几何冲刺游戏玩法中学习一个紧凑的、动作条件的世界模型。为了测试学习到的动态是否可操作,控制器通过行为克隆(BC)初始化,并完全在冻结模型回滚中使用近端策略优化(PPO)进行细化,无需与实时游戏进一步交互。在三个控制器种子上,细化后的策略在所有三个官方关卡和一个保留的社区布局上比其BC初始化存活时间更长。在部署时,基线跳过视觉生成,并在消费级GPU上维持60赫兹的捕获到动作循环。动作条件的延续和回滚诊断表明,尽管长期保真度不完美,该模型对控制仍然有用。

英文摘要

World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Dash gameplay. To test whether the learned dynamics are actionable, a controller is initialized by behavioural cloning (BC) and refined with Proximal Policy Optimization (PPO) entirely in frozen-model rollouts, without further interaction with the live game. Across three controller seeds, the refined policies survive longer than their BC initializations on all three official levels and a held-out community layout. At deployment, the baseline skips visual generation and sustains a 60-Hz capture-to-action loop on a consumer GPU. Action-conditioned continuations and rollout diagnostics show that the model remains useful for control despite imperfect long-horizon fidelity.

Comments10 pages, 2 figures. Accepted at the NeurIPS 2026 workshop "PTA: From Pretrained Representations to Acting Agents". Project page: https://tariolle.github.io/dash-vmc/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑