arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Game2World引擎:解锁野外游戏视频用于世界模型训练

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

Wenxuan Shen, Dongna Jin, Dongping Chen

arXiv 2608.24680首次发表:更新:

AI 中文总结

针对游戏视频界面干扰世界模型训练的问题,提出GameUI-Taxonomy与G2WEngine框架构建Game2World数据集,研发无掩码UI去除模型GameCleaner,提升了世界模型训练效果与UI去除性能。

AI 中文摘要

电子游戏为视频世界模型提供了可扩展的训练数据源,具备多样环境、复杂交互以及丰富的野外游戏视频。然而,原始游戏画面将游戏世界与屏幕空间界面纠缠在一起,引入了游戏特定偏差和无关动态,阻碍世界模型训练。为解决该问题,我们提出GameUI-Taxonomy与G2WEngine,这是一套用于规范化游戏UI定位与去除的全栈框架。G2WEngine可自动从真实游戏视频中提取可复用的UI资产,并在干净素材上合成时间连贯的UI叠加层。基于该引擎,我们构建了Game2World,包含96K个带有精确重建目标的合成配对视频,以及来自303款游戏的1079个野外片段用于真实评估。其资产库包含从1010个代表性游戏帧中收集的、涵盖21个分类类别的5132个经验证的UI元素。基于Game2World,我们提出GameCleaner,这是一种无掩码的游戏UI去除模型,结合了多模态语义理解与视频编辑能力。与基于掩码的方法不同,GameCleaner可直接识别并去除各类HUD元素,同时保留底层场景内容与时间动态。在受控试点中,基于无UI游戏数据训练的世界模型,其整体VideoReward相比基于带UI叠加数据训练的模型提升了6.83%。在UI去除评估中,GameCleaner在合成视频上的平均AAR达95.36,比最强的时间掩码基线高出57.3%;在野外数据上获得80.05的最佳AAR,同时实现99.8%的背景保留。这些结果证明了将互联网游戏视频转化为高质量世界模型训练数据的可扩展潜力。代码、数据集和模型将在此httpsURL发布。

英文摘要

Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the-wild gameplay videos. However, raw gameplay footage entangles the game world with screen-space interfaces, introducing game-specific biases and irrelevant dynamics that hinder world-model training. To address this problem, we introduce GameUI-Taxonomy and G2WEngine, a full-stack framework that formalizes gameplay UI grounding and removal. G2WEngine automatically extracts reusable UI assets from real gameplay videos and synthesizes temporally coherent UI overlays on clean footage. Using this engine, we construct Game2World, comprising 96K synthetic paired videos with precise reconstruction targets and 1,079 in-the-wild clips from 303 games for realistic evaluation. Its asset library contains 5,132 verified UI elements across 21 taxonomy categories, collected from 1,010 representative gameplay frames. Based on Game2World, we propose GameCleaner, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities. Unlike mask-based methods, GameCleaner directly identifies and removes diverse HUD elements while preserving the underlying scene content and temporal dynamics. In a controlled pilot, world models trained on UI-free gameplay improve overall VideoReward by 6.83% over those trained on UI-overlaid data. On UI-removal evaluation, GameCleaner achieves an average AAR of 95.36 on synthetic videos, outperforming the strongest temporal mask baseline by 57.3%, and obtains the best in-the-wild AAR of 80.05 with 99.8 background preservation. These results demonstrate the scalable potential of transforming Internet gameplay videos into high-quality world-model training data. Code, dataset, and model will be available at https://github.com/Dongping-Chen/Game2World.

CommentsWe are currently building Gaming World Model and data engine that transfers game dynamics to robotics. Feel free to contact Dongping Chen (dongpingchen0612@gmail.com) if you are interested in research collaboration or financial support

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑