arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GameEngineBench:在真实C++运行时环境中评估编码智能体

GameEngineBench: Evaluating Coding Agents on Real C++ Runtime Environments

Brian La, Sejoon Chang, Ben Kim, Junyoung Bae, Aamish Ahmad Beg, Sei Chang, Gonzalo Gonzalez-Pumariega, Kanav Goyal

arXiv 2607.03525首次发表:更新:

发表机构

Nitrode; Nexon Intelligence Labs; Dartmouth University; Columbia University; Cornell University(Nitrode; Nexon Intelligence Labs; 达特茅斯大学; 哥伦比亚大学; 康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究利用GameEngineBench在虚幻引擎5项目中评估编码智能体,核心方法是构建来自九个真实游戏仓库的基准测试集,主要贡献是揭示智能体在实时交互软件的C++开发中面临挑战,凸显游戏引擎基准测试的价值。

AI 中文摘要

游戏引擎提供实时模拟、渲染、物理、交互、网络和资产管道等功能,不仅对游戏有价值,对医疗、机器人、建筑、制造等领域的3D应用也有价值。游戏开发是这些系统最成熟且公开可用的领域,为评估编码智能体提供了实用测试平台。我们展示了GameEngineBench,这是一个用于评估虚幻引擎5项目中作用域C++实现任务的编码智能体的基准测试,它由九个真实世界游戏仓库构建而成。评估集包含110个任务,涵盖游戏玩法机制、多人行为、人工智能和世界编排、动画和移动、用户界面和会话代码、加载行为、在线服务集成、持久性、数据序列化、扩展现实行为和面向渲染的插件等。这些任务要求模型进行原生C++更改,以便在可执行的虚幻引擎项目中编译并满足行为测试。在十二个评估配置中,最强的模型达到了55.5%的pass@1,而31个任务在每个配置中都未解决。我们的结果表明,前沿编码智能体在实时交互软件的深度集成C++开发方面仍然面临困难,突出了游戏引擎基准测试作为现有软件工程评估的宝贵补充。

英文摘要

Game engines provide real-time simulation, rendering, physics, interaction, networking, and asset pipelines, making them valuable not only for games but also for 3D applications in healthcare, robotics, architecture, manufacturing, and related domains. Because game development is where these systems are most mature and publicly available, it offers a practical testbed for evaluating coding agents that must modify C++ code within stateful, interactive, real-time systems. We present GameEngineBench, a benchmark for evaluating coding agents on scoped C++ implementation tasks inside Unreal Engine 5 projects, built from nine real-world game repositories. The evaluation set consists of 110 tasks spanning gameplay mechanics, multiplayer behavior, AI and world orchestration, animation and movement, UI and session code, loading behavior, online-service integration, persistence, data serialization, XR behavior, and rendering-oriented plugins. These tasks require models to make native C++ changes that compile and satisfy behavioral tests within executable Unreal Engine projects. Across twelve evaluated configurations, the strongest model reaches 55.5\% pass@1, while 31 tasks remain unsolved by every configuration. Our results demonstrate that frontier coding agents continue to struggle with deeply integrated C++ development for real-time interactive software, highlighting game-engine benchmarks as a valuable complement to existing software engineering evaluations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑