arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GameBoyWorlds:具身视频游戏中自我改进的测试平台

GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games

Dhananjay Ashok, Adam Shen, Aslan Huo Feng, Chinmay Khanna, Jun Rui Huang, Raghav Sarmukaddam, Surendira Balaji Natarajan, Xiaotong Cui, Xincan Zhang, Thomson Yen, Hongseok Namkoong, Jonathan May, Jesse Thomason

arXiv 2609.32093首次发表:更新:

发表机构

Information Sciences Institute, University of Southern California; Columbia University; University of Chicago; Purdue University; Georgia Institute of Technology(南加州大学信息科学研究所; 哥伦比亚大学; 芝加哥大学; 普渡大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GameBoyWorlds是一个用于评估具身视频游戏中智能体自我改进能力的测试平台,通过执行和通关任务证明前沿模型因多模态接地失败而表现不佳,现有自我改进方法亦不足,需从自身经验中自主提升。

AI 中文摘要

在专家指导下,智能体能够在交互式环境中运行;然而,它们能否从自身经验中自主学习的尚不清楚。为评估此类自我改进方法,我们引入了GameBoyWorlds,一个用于视频游戏中智能体自我改进的测试平台。GameBoyWorlds-Execution在5个不同游戏系列的集合上评估任务执行能力。智能体被允许访问专门的训练游戏,但未提供演示、文档或奖励。智能体必须通过自主探索并在环境中自我定位,并从自身经验中推断可操作的知识。在测试时,智能体必须完成短视界任务,这些任务评估其在未见过的游戏中导航、交互和参与游戏特定机制的能力。开箱即用的前沿模型因多模态接地失败而未能完成500个任务中的50%以上,这表明自我改进方法仍有提升性能的空间。我们证明了当代自我改进方法存在不足,世界建模和自主技能发现失败,而一种使用基于好奇心的探索来编写指南的新策略仅取得部分成功。GameBoyWorlds-Playthrough测试两款粉丝制作的宝可梦游戏中的端到端游戏完成情况。我们表明,尽管前沿模型已预先接触过官方发布版本(如宝可梦红),但它们缺乏关于我们测试平台中游戏的基本信息。智能体必须从自身经验中学习,并在游戏过程中自主改进,而不是依赖其参数化知识来取得成功。我们展示了一个具有多模态记忆和层次化子目标的复杂智能体流程在两款游戏中均未能达到第一个主要里程碑,从而将GameBoyWorlds确立为自我改进智能体的一个雄心勃勃的目标。

英文摘要

Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. At test time, agents must complete short-horizon tasks that evaluate their ability to navigate, interact, and engage with game-specific mechanics in unseen games. Out-of-the-box frontier models complete fewer than 50% of the 500 tasks due to failures in multimodal grounding, establishing that self-improvement methods have room to push performance. We demonstrate that contemporary approaches to self-improvement are lacking, with world modelling and autonomous skill discovery failing, and a novel strategy that uses curiosity-based exploration to write guides achieving only partial success. GameBoyWorlds-Playthrough tests end-to-end game completion in two fan-made Pokémon games. We show that while frontier models have been pre-exposed to official releases such as Pokémon Red, they lack essential information on the games in our testbed. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough. We show that a sophisticated agentic pipeline with multimodal memory and hierarchical subgoals fails to reach even the first major milestone in both games, establishing GameBoyWorlds as an ambitious target for self-improving agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑