递归游戏创建器:一种面向产品级体验的游戏智能体开发框架
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness
- HKU MMLab(香港大学多媒体实验室)
- The University of Hong Kong(香港大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出递归游戏创建器,通过设计者、构建者、玩家和评审者四组件递归协作,将原型游戏优化为娱乐性产品,在GameCraft-Bench和GameASG-Bench上取得领先性能。
AI中文摘要:
近期游戏设计智能体在生成可玩游戏方面取得了显著进展。然而,程序的正确性并不能保证玩家获得愉悦的体验。我们提出了递归游戏创建器(Recursive Game Creator),一种面向体验的框架,旨在将智能体游戏开发从粗糙的游戏原型推进到具有娱乐性的成品游戏。递归游戏创建器围绕四个组件组织递归开发:设计者(Designer)、构建者(Builder)、玩家(Player)和评审者(Reviewer)。设计者将用户指令和评审者的反馈转化为详细的计划。构建者将这些计划转化为候选游戏。编码原生的玩家通过程序化接口创建并执行可复用的策略,以高效收集多样化的游戏轨迹,从而缓解因基于GUI的缓慢收集而导致的评估偏差。评审者使用精心设计的基于轨迹的指标来诱导玩家偏好,结合视觉证据和明确的文本偏好,根据游戏特定标准对游戏进行评估。最后,评审者接受更优版本并为下一轮提供改进意见,从而闭合递归循环。我们的方法在GameCraft-Bench上达到了最先进的整体性能77.89。在GameASG-Bench上,它实现了53.2%的严格任务成功率,比同模型基线提高了34.1%,并且在比较方法中平均运行时检查通过率最高,达到93.4%。一项用户研究显示更长的游玩时间和更高的评分。代码即将发布。
英文摘要:
Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.