发表机构
Institute of Automation, Chinese Academy of Sciences; Baidu Inc.(中国科学院自动化研究所; 百度公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GameGo框架,将简短游戏种子转化为产品需求文档,通过动态压缩训练编码智能体,构建55,060条轨迹数据集和124查询基准,使模型性能超越基线并接近前沿。
AI 中文摘要
近年来,大型语言模型(LLM)的进展展示了其在Web前端执行方面的显著能力,其中基于浏览器的游戏生成已成为一个尤为突出的前沿领域。尽管以往的工作常常依赖于复杂的多轮工作流或专注于静态游戏评估基准,但本研究旨在通过编码智能体直接实现端到端的真实世界游戏合成。然而,直接从稀疏的用户查询生成复杂游戏往往迫使编码智能体做出不明确的假设,导致机制不完整、游戏流程脱节以及视觉美感有限。为解决这一问题,本文提出了GameGo,一个可扩展的框架,系统性地将简短的游戏种子转化为基于行业游戏开发实践的全面产品需求文档。为了在保留核心游戏机制约束的同时不限制设计探索,GameGo采用任务特定的动态压缩,以最大化信息密度并保持指令遵循。基于此流程,构建了GameGoData,包含55,060条跨越2D、2.5D和3D游戏的开发轨迹,以及GameGoBench,一个包含124个多样化游戏查询的基准。在GameGoData上训练GameGoCoder得到的模型,在游戏开发基准上优于匹配的基线模型,并与前沿模型相当。所有代码、数据集和模型将公开发布。
英文摘要
Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities in web front-end execution, with browser-based game generation emerging as a particularly prominent frontier. While previous efforts frequently rely on complex multi-turn workflows or focus on static game evaluation benchmarks, this work targets direct end-to-end real-world game synthesis driven by coding agents. However, generating complex games directly from sparse user queries often forces coding agents to make underspecified assumptions, yielding incomplete mechanics, disconnected gameplay flows, and limited visual aesthetics. To resolve this issue, this paper presents GameGo, a scalable framework that systematically transforms brief game seeds into comprehensive Product Requirements Documents grounded in industry game-development practices. To retain core gameplay constraints without restricting design exploration, GameGo uses task-specific dynamic compression to maximize information density while preserving instruction following. Based on this pipeline, GameGoData is constructed with 55,060 development trajectories across 2D, 2.5D, and 3D games, alongside GameGoBench, a benchmark comprising 124 diverse game queries. Training GameGoCoder on GameGoData yields a model that outperforms matched baselines and is comparable to frontier models across gamedev benchmarks. All code, datasets, and models will be made publicly available.