arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RSIGame:具有递归自我改进的自主智能体游戏开发

RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement

Wenyi Wu, Minghao Fu, Jieyu You, Kun Zhou, Siqi Liu, Aayush Salvi, Yiheng Lin, Ce Zhang, Xiaohan Lan, Jiahui Zhu, Yujie Zhong, Qi She, Biwei Huang

arXiv 2609.39045首次发表:更新:

发表机构

University of California San Diego; ByteDance Inc.; Carnegie Mellon University(加州大学圣迭戈分校; 字节跳动有限公司; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RSIGame通过局部探索-诊断-改进循环和全局质量跟踪循环,实现自主游戏开发的递归自我改进,并在140个任务上显著提升游戏质量,同时将经验内化至生成器,使小模型超越大模型并大幅降低生成成本。

AI 中文摘要

大型语言模型的最新进展使得自动游戏生成日益可行,然而,在可玩版本之外可靠地改进生成的游戏仍然具有挑战性。朴素的迭代细化容易过度拟合一小部分测试用例,产生脆弱的游戏,存在未解决的错误、缺失的行为,并且对更广泛的玩家交互泛化能力差。我们引入了RSIGame,一个具有递归自我改进的自主智能体游戏开发框架。RSIGame将开发组织为互补的局部和全局循环。具体而言,局部探索-诊断-改进循环广泛探索可执行游戏,诊断并优先处理发现的问题,并进行基于证据的修订,其中不断演进的检查清单持续积累新的测试和改进指导。全局循环跟踪整体质量,保留最佳检查点,并检测长期开发中的饱和或回归。除了测试时改进,RSIGame还通过训练将成功的开发经验内化到生成器中。在140个GameCraft-Bench任务、两个游戏引擎和五个生成器中,RSIGame在匹配的开发预算下持续提高游戏质量。值得注意的是,经验内化使Qwen3.8-27B在Godot上达到61.38,在Phaser上达到58.53,超过了GPT-5.5的一次性得分,同时将Qwen的生成令牌减少了11倍。

英文摘要

Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑