SWE-Game:编码智能体能否构建我们想要的游戏?
SWE-Game: Can Coding Agents Build the Games We Want?
- Shanghai Jiao Tong University(上海交通大学)
- Zhongguancun Academy(中关村学院)
- Shanghai Innovation Institute(上海创新研究院)
- Shenzhen University(深圳大学)
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Elbetech Technology(埃尔贝特科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SWE-Game基准测试评估编码智能体在游戏开发中的能力,涵盖247个任务,结果显示最佳模型得分仍低于60,主要问题为需求遗漏和逻辑错误,并验证了运行时检查与视觉评估结合的有效性。
AI中文摘要:
我们介绍了SWE-Game,这是一个包含247个任务的基准测试,这些任务基于41个可执行的参考Godot游戏,涵盖2D和3D中的13个游戏玩法类别。五种任务类型涵盖了从简要说明开始的开发、根据游戏设计文档的实现、骨架补全、修复83个注入故障案例,以及Godot到Unity的移植。参考材料指定了预期的游戏玩法,而共享的仪器接口允许评估者拥有的驱动器和探针执行操作并观察独立实现的游戏。评估结合了引擎状态检查、认证参考输入重放和智能体撰写的功能演示,以评估机制正确性、展示的可玩性以及修复后的行为恢复和保留。特定于游戏的视觉语言评分标准分别评估呈现效果。在六个模型中,Opus5在所有五种任务类型中取得了最高的总体得分。在三个构建任务中,最佳总体得分仍低于60分(满分100分),其中Brief-to-Game达到50.38分。对已审阅提交的分析确定了需求遗漏和游戏逻辑错误是主要的实现问题。在来自100个智能体构建游戏的人工标记行为上,可执行检查实现了92.59%的平衡准确率,而基于视频的VLM评判器为78.41%。基于评分标准的视觉得分与200个游戏片段的人工评分达到了0.829的Spearman相关系数。这些结果共同刻画了当前智能体在游戏开发活动中的能力,并支持将运行时证据与视觉评估相结合。
英文摘要:
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemented games. Evaluation combines engine-state checks, certified reference-input replay, and agent-authored feature demonstrations to assess mechanic correctness, demonstrated playability, and behavioral restoration and preservation after repairs. Game-specific vision-language rubrics separately assess presentation. Across six models, Opus5 achieves the highest overall score in all five task types. Best overall scores remain below 60 out of 100 across the three construction tasks, with Brief-to-Game reaching 50.38. Analysis of reviewed submissions identifies requirement omissions and gameplay logic errors as predominant implementation problems. On human-labeled behaviors from 100 agent-built games, executable checks achieve 92.59% balanced accuracy, compared with 78.41% for a video-based VLM judge. Rubric-based visual scores reach a Spearman correlation of 0.829 with human ratings of 200 gameplay clips. Together, these results characterize current agent capabilities across game-development activities and support combining runtime evidence with visual assessment.