arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

A2Z GameSpec-Bench:编码智能体从游戏设计规格生成游戏的忠实度如何?

A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?

Seonho Lee, Wonryeol Jeong, Alberto Cereser, Inha Kang, Hyeonjong Kim, Seungmin Kwak, Dongmin Park

arXiv 2609.39564首次发表:更新:

发表机构

KRAFTON; KAIST; Korea National University of Arts(魁匠团; 韩国科学技术院; 韩国艺术综合大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出A2Z GameSpec-Bench基准,含100个长篇GDD,通过依赖感知契约和测试策略评估编码智能体生成游戏的忠实度,发现现有智能体难以满足相互依赖需求,特定需求反馈可提升10.9%忠实度。

AI 中文摘要

将完整的应用程序开发委托给编码智能体,要求保留预期的设计,而非通过简单的提示生成看似合理的输出。游戏开发提供了一个要求严格的测试平台,因为长篇游戏设计文档(GDD)描述了必须在游戏逻辑、视觉渲染和玩家交互中协同工作的需求。然而,现有的游戏开发基准通常使用紧凑的规格,并且对评估长篇GDD中这些方面相互依赖的需求支持有限。我们引入了A2Z GameSpec-Bench,一个包含100个长篇GDD的基准,用于评估智能体的端到端游戏开发。我们通过检查游戏是否满足GDD需求并保留它们之间的关系来衡量忠实度。每个GDD被转化为一个依赖感知的契约,包含规则、约束和先决条件关系。遵循游戏开发实践,我们将源代码检查与智能体生成的测试策略相结合,用于基于场景的回放和自适应游戏测试。该契约在智能体和修订轮次之间保持固定,而与相同需求关联的判断和证据支持一致的比较和失败检测。我们的评估表明,当前智能体难以在代码实现和实际游戏中共同满足相互依赖的需求。特定需求的反馈在两轮后相对于自我修订将GDD忠实度提高了10.9%。A2Z GameSpec-Bench评估了超越实现判断的端到端规格遵循能力,并提供了有针对性的反馈以支持更忠实的游戏开发。代码和数据集可在该https URL获取。

英文摘要

Delegating complete application development to coding agents requires preserving the intended design rather than simply producing plausible outputs through naive prompting. Game development provides a demanding testbed, as long-form Game Design Documents (GDDs) describe requirements that must work together across game logic, visual rendering, and player interactions. However, existing game-development benchmarks typically use compact specifications and provide limited support for evaluating interdependent requirements across these aspects in long-form GDDs. We introduce A2Z GameSpec-Bench, a benchmark of 100 long-form GDDs for evaluating end-to-end game development by agents. We measure faithfulness by checking whether the game satisfies the GDD requirements and preserves the relationships among them. Each GDD is turned into a dependency-aware contract that contains rules, constraints, and prerequisite relations. Following game-development practices, we combine source-code inspection with agent-generated test policies for scenario-based replay and adaptive playtesting. The contract remains fixed across agents and revision rounds, while judgments and evidence linked to the same requirements support consistent comparison and failure detection. Our evaluations show that current agents struggle to jointly satisfy interdependent requirements across code implementation and actual play. Requirement-specific feedback improves GDD Fidelity by 10.9% relative to self-revision after two rounds. A2Z GameSpec-Bench assesses end-to-end specification-following ability beyond implementation judgments and provides targeted feedback to support more faithful game development. Code and datasets are available at https://a2z-gamespec-bench.github.io.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑