GameLogicBench:基于逐帧状态断言的运行时游戏逻辑编码智能体评估
GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有游戏开发基准缺乏运行时规则检查的问题,提出GameLogicBench基准,通过逐帧状态断言和变异体验证,在72个任务上评估编码智能体,发现其性能随任务范围扩大而下降,且可靠评估需同时考虑测试拒绝和外部代码访问。
AI中文摘要:
编码智能体能够跨大型软件项目修改和测试代码。游戏开发是智能体必须实现游戏规则的领域。游戏即使在运行过程中违反规则,也可能以有效状态结束。当前游戏开发基准要么回放固定示例、对视频评分,要么让另一个模型评判结果。然而,现有基准均未在多种评估者选择的场景中检查整个执行过程中的游戏规则,同时确保完全可复现的判定。我们提出了GameLogicBench,一个包含72个Godot项目中游戏逻辑任务的基准。一个自动化评估器在每个模拟帧检查每个游戏的规则。在403个手工设计的场景中,种子参数变化产生了1,451个测试用例。为确保评估器衡量行为而非实现选择,它必须接受每个任务的多种正确实现,同时拒绝变异体(即移除了一项所需能力的实现)。这些任务涵盖孤立机制、多系统交互和仓库级功能。在20种语言模型与脚手架的组合中,最佳观察运行解决了52.78%的任务。在Claude Code下,随着任务范围从孤立机制扩展到交互系统再到仓库级功能,所有十二个模型解决的任务数减少。智能体在仓库级任务上比在孤立机制任务上更频繁地检查代码并进行更多工具调用。大多数不成功的提交是可运行的,但错误地实现了某些所需的游戏行为。我们比较了使用和不使用变异体验证的基准评估器版本。没有此验证,错误的智能体提交会通过。一项单独的分析发现,当网络访问开放时,智能体会从公共仓库复制代码。因此,可靠的评估既取决于测试拒绝什么,也取决于外部代码智能体可以访问什么。
英文摘要:
Coding agents can modify and test code across large software projects. Game development is a domain where agents must implement gameplay rules. A game can end in a valid state even after violating its rules during the run. Current game-development benchmarks replay fixed examples, score videos, or ask another model to judge the result. However, no existing benchmark checks game rules throughout execution across varied evaluator-selected scenarios while ensuring exactly reproducible verdicts. We introduce GameLogicBench, a benchmark of 72 gameplay-logic tasks in Godot projects. An automated evaluator checks each game's rules at every simulation tick. Across 403 hand-designed scenarios, seeded parameter variations produce 1,451 test cases. To ensure that the evaluator measures behavior rather than implementation choice, it must accept different correct implementations for each task while rejecting mutants, implementations with one required capability removed. The tasks span isolated mechanics, multi-system interactions, and repository-scale features. Across 20 combinations of language models and scaffolds, the best observed run solves 52.78% of tasks. Under Claude Code, all twelve models solve fewer tasks as task scope expands from isolated mechanics, through interacting systems, to repository-scale features. Agents inspect code more often and make more tool calls on repository-scale tasks than on isolated-mechanic tasks. Most unsuccessful submissions are runnable, but implement some required game behavior incorrectly. We compared versions of our benchmark evaluator built with and without validation using mutants. Without this validation, incorrect agent submissions passed. A separate analysis finds agents copying code from public repositories when network access is open. Reliable evaluation thus depends both on what the tests reject and on what external code agents can access.