持续Harness:面向自我改进基础代理的在线适应
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
- Princeton University(普林斯顿大学)
- ARISE Foundation(ARISE基金会)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出持续Harness,一种无需人工干预的自我改进机制,通过在线适应提升具身代理在长视界部分可观察决策中的表现,实验证明其在Pokémon游戏中显著降低操作成本并接近专家水平。
AI中文摘要:
编码Harness如Claude Code和OpenHands通过工具、记忆和规划包装基础模型,但尚未有类似方法用于具身代理的长视界部分可观察决策。我们首先报告了Gemini Plays Pokemon(GPP)实验。通过迭代的人工在环Harness优化,GPP成为首个完成Pokémon Blue、Yellow Legacy和Crystal的AI系统。在最艰难阶段,代理自身通过长上下文记忆迭代策略,产生自我改进信号。持续Harness移除了人工干预:一种无重置的自我改进Harness,正式化并自动化了所观察到的现象。从仅有的最小环境接口开始,代理交替执行和优化自身提示、子代理、技能和记忆,利用任何过去轨迹数据。提示优化方法需要回合重置;持续Harness在单次运行中在线适应。在Pokémon红和Emerald上,持续Harness从头开始显著降低按钮按压成本,相比最小化基线并恢复大部分与手工程专家Harness的差距,尽管从相同的原始接口开始,无定制知识、无手工工具和无领域支撑。随后通过模型自身闭合循环:一个在线过程-奖励共学习循环,其中开源代理通过优化Harness的rollouts被手稿模型重新标记并用于更新模型,驱动Pokémon红中的持续游戏里程碑进展,无需在训练迭代之间重置环境。
英文摘要:
Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agents' long-horizon partial-observability decision-making. We first report our Gemini Plays Pokemon (GPP) experiments. With iterative human-in-the-loop harness refinement, GPP became the first AI system to complete Pokemon Blue, Yellow Legacy on hard mode, and Crystal without a lost battle. In the hardest stages, the agent itself began iterating on its strategy through long-context memory, surfacing emergent self-improvement signals alongside human-in-the-loop refinement. Continual Harness removes the human fully from this loop: a reset-free self-improving harness for embodied agents that formalizes and automates what we observed. Starting from only a minimal environment interface, the agent alternates between acting and refining its own prompt, sub-agents, skills, and memory, drawing on any past trajectory data. Prompt-optimization methods require episode resets; Continual Harness adapts online within a single run. On Pokemon Red and Emerald across frontier models, Continual Harness starting from scratch substantially reduces button-press cost relative to the minimalist baseline and recovers a majority of the gap to a hand-engineered expert harness, with capability-dependent gains, despite starting from the same raw interface with no curated knowledge, no hand-crafted tools, and no domain scaffolding. We then close the loop with the model itself: an online process-reward co-learning loop, in which an open-source agent's rollouts through the refining harness are relabeled by a frontier teacher and used to update the model, drives sustained in-game milestone progress on Pokemon Red without resetting the environment between training iterations.