arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21293cs.AIcs.SE

GameASG-Bench:面向游戏开发的自主软件生成基准测试

GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development

  • Ant Group(蚂蚁集团)
  • Beihang University(北京航空航天大学)
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiuhui Zhang, Yi Chen, Shusheng Xu, Fan Li, Huan Wang, Tongkai Yang, Binhang Yuan

AI总结:

GameASG-Bench提出面向游戏开发的行为可测试自主软件生成基准,通过L1静态与L2浏览器执行检查,评估九个智能体栈,揭示高平均通过率掩盖的任务级合规差距。

AI中文摘要:

自主软件生成(ASG)旨在将人类需求转化为可执行的应用程序,但交付这些应用程序并不必然意味着其交互组件满足指定的行为需求。我们引入了GameASG-Bench,一个将行为可测试性纳入游戏开发生成任务的基准。我们的设计在生成之前声明评估接口规范,固定合法的起始场景、玩家级动作、稳定快照、拒绝行为和不变量,同时保留私有实现的开放性。具体而言,我们包括:(i)静态L1检查,评估源代码级合规性;以及(ii)浏览器执行的L2检查,将语义观察与真实输入和运行时证据相结合。我们将此协议实现为47个浏览器原生游戏生成任务,涵盖12个主要类型以及2D和3D交互,每个任务都有可执行的检查和独立验证的参考实现。我们的实验回答了关于端到端智能体性能、工具访问和名义回合预算、推理努力以及测试框架选择的四个关键问题。在九个智能体栈中,观察到的最高平均L2检查通过率为93.2%,但观察到的最高严格任务成功率(要求所有L1和适用的L2先决条件及核心需求检查)仅为55.3%(26/47个任务)。对于DeepSeek-V4-Flash,完整的工具访问和更大的名义回合预算带来更多的严格任务成功,而严格任务成功率在推理努力方面不是单调的。两个测试框架均实现了18个严格任务成功,但只有十个任务在两个框架下都成功。这些结果暴露了任务级合规性差距,而高平均检查通过率实际上掩盖了这些差距。

英文摘要:

Autonomous software generation (ASG) aims to turn human requirements into executable applications, but delivering these applications does not necessarily establish that their interacting components satisfy the specified behavioral requirements. We introduce GameASG-Bench, a benchmark that makes behavioral testability part of the generation task for game development. Our design declares an evaluation interface specification before generation, fixing legal starting scenarios, player-level actions, stable snapshots, rejection behavior, and invariants while leaving private implementations open. Concretely, we include: (i) static L1 checks that assess source-level compliance; and (ii) browser-executed L2 checks that combine semantic observations with real input and runtime evidence. We implement this protocol as 47 browser-native game-generation tasks spanning 12 primary genres and both 2D and 3D interaction, each with executable checks and an independently verified reference implementation. Our experiments answer four key questions about end-to-end agent performance, tool access and nominal turn budget, reasoning effort, and harness choice. Across nine agent stacks, the highest observed mean L2 check pass rate is 93.2%, yet the highest observed strict task success rate, requiring all L1 and applicable L2 prerequisite and core requirement checks, is only 55.3% (26/47 tasks). For DeepSeek-V4-Flash, full tool access and larger nominal turn budgets yield more strict task successes, while the strict task success rate is not monotonic in reasoning effort. Both tested harnesses achieve 18 strict task successes, but only ten tasks succeed under both. These results expose task-level compliance gaps that high average check pass rates actually obscure.

补充信息

↑