发表机构
New York University; Princeton University(纽约大学; 普林斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Gauntlet框架,验证前沿通用编码智能体能在无辅助情况下,通过单次自主会话从裸交互构建完整游戏控制器,首次独立赢得星际争霸II和文明等完整商业游戏,展现编译智能体能力。
AI 中文摘要
LLM智能体在将游戏知识转化为熟练游戏表现方面屡屡受挫,即使研究人员围绕模型构建智能体——提供感知、记忆、技能库、规划器或可执行策略脚手架。编码智能体的快速进步提出了两个更尖锐的问题:前沿模型现在能否赢得游戏,以及它们能否在无辅助的情况下,完全自行构建整个玩家?我们引入Gauntlet,一个开发-冻结-评估框架,将游戏从小型街机移植到完整的商业规模作品,背后是一个刻意简化的契约:通用编码智能体接收游戏描述、原始观察/动作接口和一个空策略文件——没有策略、没有算法、没有架构。在单次自主会话中,智能体与实时游戏交互并工程化一个独立控制器;我们冻结结果并在保留实例上评分,游戏过程中零模型调用。在一个未公开的程序化Roguelike游戏上,保留实例的成功率跨度从0%到86%,并暴露了一个尖锐的代际阈值:新一代系统的每个观察会话都优于其前代的最佳会话。在全游戏规模上,一个编译的原始API控制器击败了所有公平的星际争霸II内置AI和两个作弊变体,单会话程序通过完全征服在保留种子上赢得了完整的文明(Freeciv)游戏。尽管对新手AI的胜率不高,但这是首次:此前没有语言智能体系统能在无逐回合模型调用和手工战术层的情况下,独立赢得该类型的完整游戏。前沿编码智能体开始追踪长时程策略。冻结的程序是可检查的。我们将这种能力称为编译智能体:开发经验被编译成一个持久的可执行智能体,其架构由模型构建。
英文摘要
Coding agents are increasingly capable of sustained engineering and empirical research. Greater autonomy makes their research decisions themselves a target for evaluation: what to investigate, which experiments to run, and when to stop. We introduce Gauntlet, a develop-freeze-evaluate protocol for studying these decisions as agents build standalone game-playing programs. Starting from a game description, a raw observation/action interface, and an empty policy file, an off-the-shelf agent develops a controller from bare interaction in one autonomous session. We call the capability under study compiled agency: the shipped program plays with zero model calls. The capability is real and advancing: environment access adds 10 to 78 percentage points of held-out success over construction-only controls, and progress across model generations comes in steps, with tiers that defeat one generation entirely falling to the next. The protocol yields StarCraft II controllers that defeat every fair built-in AI, and in the newest generation the strongest cheating tier as well, and Civilization controllers that win complete games by conquest. Replaying more than 5,000 frozen versions exposes the research behind the programs: gains that plateau early; rigorous local investigation beside sparse validation of what actually ships; and stops that follow a race between the agent's own evidence and a model-specific transcript budget it was never given. When validation panels are refreshed mid-session, so that only the evidence changes, shipped success rises 12.4 points in nine of nine completed pairs of twelve initiated. The experiments an agent designs are part of the capability it delivers; Gauntlet makes them measurable and improvable--a step toward agents whose research practice, not just whose code, can be engineered. We release the benchmark, the corpus, and the replay tooling.