发表机构
Institute of Automation, Chinese Academy of Sciences; Tsinghua University; The Hong Kong University of Science and Technology; The Chinese University of Hong Kong, Shenzhen; Lightspeed Studios, Tencent(中国科学院自动化研究所; 清华大学; 香港科技大学; 香港中文大学(深圳); 腾讯光速工作室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GameXpert-Bench是覆盖游戏开发全生命周期的基准,含三个赛道,测评显示当前编码智能体在生成可玩游戏基础上表现较好,在缺陷发现等方面仍有不足。
AI 中文摘要
近期的大型语言模型(LLM)可作为编码智能体,根据自然语言请求构建完整游戏。游戏开发要求极高,程序逻辑、视觉与音频内容、界面、交互性及可玩性必须共同构成可执行产物,因此评估该能力需同时考量游戏产物与开发过程。现有基准通常仅通过评估最终产物或孤立开发阶段来测评LLM的游戏开发能力。我们对人类与智能体的完整开发轨迹分析后,确定了编码智能体游戏开发生命周期的三个阶段:初始游戏生成、错误诊断与修复、多轮优化。为此,我们推出GameXpert-Bench,将这三个生命周期阶段转化为三个互补的基准赛道:GameGen评估在空工作区中根据单一请求完成完整游戏创建;GameFix评估在缺陷被报告或由智能体自行发现时的诊断与修复能力;GameOpt评估基于真实用户与智能体开发轨迹生成的请求链实现的累积优化能力。我们通过实时游戏交互、确定性行为测试或带有回归检查的最终产物标准来评估每个赛道。该基准套件包含11个类型的97个生成任务;50个游戏关卡经人工验证的100个修复任务,每个任务注入19-27个错误;以及17个优化链,含6轮交互与102个请求。在三个赛道中,当前智能体在生成可玩基础、实现明确需求方面更可靠,但在发现缺陷、验证运行时行为及在变更中保留功能方面表现欠佳。
英文摘要
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.