SpeedrunBench:挑战LLM代理的视频游戏速通
SpeedrunBench: Challenging LLM Agents with Video Game Speedrunning
浏览论文内容
中文总结 AI 辅助
本文提出SPEEDRUNBENCH基准,通过9款视频游戏速通任务评估LLM代理的策略形成能力,实验显示代理在简单游戏中接近人类纪录,但在复杂游戏中仍落后。
中文摘要 AI 辅助
前沿LLM代理已被证明能够解决人类有可衡量解决方案的日益复杂的任务。这引发了一个相关问题:LLM代理能否超越人类已解决的问题。当人类开发的成熟解决方案对于缺乏背景或足够训练数据的问题变得不足时,开发复杂策略以应对重大问题能力变得至关重要。我们通过视频游戏速通的社区实践来研究代理的这种策略形成能力。在速通中,实践者竞争在特定条件下找到完成视频游戏的最快方式,并在此过程中发现需要深入理解和掌握底层游戏机制的非正统玩法。我们引入了SPEEDRUNBENCH,一个评估前沿LLM代理在9个不同游戏上的基准。为了在此基准中表现出色,代理必须反复改进其策略,反思其表现,利用其获得的知识,并在长时间跨度的行动中进行推理,以改进一个日益困难的问题:比自己和他人更快。我们的实验表明,虽然前沿代理在简单平台游戏中接近人类世界纪录,但在实际预算下,它们在更长、更复杂的游戏中仍落后于人类表现。这些结果表明,SPEEDRUNBENCH是研究代理策略形成能力的有用测试平台,也是一种抗饱和的评估措施,因为几乎总有更快的完成时间等待被发现。
英文摘要
Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.
发表机构
- Patronus AI
- University of Washington(华盛顿大学)
- Institute of Science Tokyo(东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。