arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActSWM:面向开放世界游戏长视界规划的动作敏感型世界模型

ActSWM: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games

Zhenfeng Gan, ZiTong Zeng, Jiajun Cheng, Yeke Song, Yongyi Tang, Xueqian Wang

arXiv 2607.26712首次发表:更新:

AI 中文总结

针对开放世界游戏长视界规划的上下文崩塌问题,提出动作敏感型世界模型ActSWM,通过约束隐式回滚保留动作依赖差异,提升了任务成功率与动作恢复能力。

AI 中文摘要

隐式世界模型通过在隐空间中优化未来控制序列并以后退视界方式重新规划,支持高效的模型预测控制。然而,现有隐式预测器往往缺乏稳定的长视界回滚能力,仅预测精度无法保证回滚过程对所规划动作保持响应性。我们识别出“上下文崩塌”这一失效模式:自回归隐式预测器在不同动作序列下会产生几乎无差异的未来状态,同时与未来状态保持高度相似性。为解决该问题,我们提出ActSWM,一种基于过渡分离原则的动作敏感型隐式世界模型:对规划有用的隐式动力学模型应保持不同动作对应的未来状态可区分,且每个局部过渡关联的动作可恢复。基于此原则,动作敏感性被作为隐式回滚的约束条件,而非仅作为辅助预测目标,从而鼓励预测的未来状态在长视界范围内保留动作依赖差异。通过步长漂移分析、闭环Minecraft规划及跨游戏局部动作恢复实验,ActSWM相比现有基线方法能保持更大的动作依赖回滚间隙,在长视界交互场景中提升任务成功率,还能基于离线游戏视频实现世界模型驱动的动作恢复。

英文摘要

Latent world models support efficient model-predictive control by optimizing future control sequences in latent space and replanning in a receding-horizon manner. However, existing latent predictors often lack stable long-horizon rollout ability, and prediction accuracy alone does not ensure that rollouts remain responsive to the actions being planned. We identify Context Collapse, a failure mode in which autoregressive latent predictors maintain high similarity to future states while producing nearly indistinguishable futures under different action sequences. To address this issue, we propose ActSWM, an action-sensitive latent world model grounded in a transition-separation principle: a planning-useful latent dynamics model should keep alternative-action futures distinguishable and make the action associated with each local transition recoverable. Under this principle, action sensitivity is enforced as a constraint on latent rollouts rather than treated only as an auxiliary prediction target, encouraging predicted futures to preserve action-dependent differences over long horizons. Across step-drift analysis, closed-loop Minecraft planning, and cross-game local action recovery, ActSWM preserves larger action-dependent rollout gaps than existing baselines, improves task success in long-horizon interactive settings, and enables world-model-based action recovery from offline gameplay videos.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑