发表机构
Roblox(Roblox)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
OpenGameEval在Roblox Studio中评测智能体编程,通过分离观察与编辑工具直接衡量探索,发现探索行为显著预测任务成功率,且当前模型在84个任务上表现困难。
AI 中文摘要
我们提出了OpenGameEval,一个用于Roblox Studio内智能体游戏开发的基准测试和评估框架。它在可复现的、有状态的游戏引擎会话中将语言模型作为智能体运行,并通过可执行的检查对每次运行进行评分,检查既针对编辑后的场景,也针对模拟的游戏会话。大多数智能体编程基准测试要求探索,但仅对最终任务成功进行评分。OpenGameEval在其八工具动作空间中分离了观察工具与编辑工具,因此可以直接衡量探索。我们测量了13个前沿模型在84个人工策划的核心任务上的通过率和探索行为,每个任务进行16次尝试。这些任务对当前模型而言是困难的。最佳模型在单次尝试中解决了51.7%的任务,在五次尝试中五次全部成功解决了39.4%的任务,且没有测试模型解决了六个任务。前沿模型通过解决不同任务达到相似的通过率:按任务所需的工作类型划分,前五名在脚本编写任务上相差5.0个百分点,在场景更改任务上相差12.5个百分点。探索行为可以预测一次运行是否成功。在固定任务和模型的情况下,在仅场景任务中,对参考解决方案接触的每个对象在操作前进行检查的运行,比不检查任何对象的运行通过率高13.4个百分点;在仅脚本任务中,这一数字为9.8个百分点。我们在MIT许可下发布了任务套件、其位置文件、每任务注释、一个在Roblox Studio内运行任务的插件,以及更新的排行榜,网址为https://this https URL。
英文摘要
We present OpenGameEval, a benchmark and evaluation framework for agentic game development inside Roblox Studio. It runs language models as agents in reproducible, stateful game-engine sessions and scores each run with executable checks, both on the edited scene and in a simulated play session. Most agentic coding benchmarks require exploration but score only final task success. OpenGameEval separates observation tools from editing tools in its eight-tool action space, so exploration can be measured directly. We measure the pass rates and exploration behavior of 13 frontier models on 84 human-curated core tasks, with 16 attempts per task. The tasks are hard for current models. The best model solves 51.7% of tasks on a single attempt and 39.4% five times out of five, and no tested model solves six of the tasks. Models at the frontier reach similar pass rates by solving different tasks: splitting tasks by the kind of work they require spreads the top five by 5.0pp on script-authoring tasks and 12.5pp on scene-change tasks. Exploration behavior predicts whether a run succeeds. Holding task and model fixed, a run that inspects every object a reference solution touches before acting on it passes 13.4pp more often than a run that inspects none of them on scene-only tasks, and 9.8pp more often on script-only tasks. We release the task suite, its place files, the per-task annotations, a plugin that runs the tasks inside Roblox Studio, and an updated leaderboard under the MIT license at https://github.com/Roblox/open-game-eval.
CommentsA shorter version appears at the NeurIPS 2026 Workshop on Evaluation of Interactive Agents