发表机构
University of Illinois Urbana–Champaign; Meta AI(伊利诺伊大学厄巴纳-香槟分校; Meta AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 EvoHarness-RL 方法,通过学习管控器策略解决长视野 LLM 智能体的外部状态管理问题,在 ALFWorld 上达到 96.9% 的任务成功率,验证了可训练管控器策略的有效性。
AI 中文摘要
长视野大语言模型(LLM)智能体越来越依赖外部执行支持来维持状态、跟踪进度、调用工具、验证结果并在交互间复用经验。然而,有效的管控器使用会带来两个相互关联的挑战:从有噪声的交互轨迹中形成状态,以及对外部状态访问进行运行时控制。现有智能体通常通过提示、启发式方法或特定领域约定来处理这两个问题,导致外部工作空间及其使用策略需人工设计。为解决该问题,我们研究管控器策略学习问题,即智能体离线学习管控器策略,并在运行时任务执行期间部署这些策略以构建和更新外部管控器状态。我们提出 EvoHarness-RL,它将信念(Belief)、进度(Progress)和经验(Experience,BPE)作为面向策略的管控器状态。有监督的管控器微调可让基础智能体掌握管控器动作空间及如何构建有用的外部状态,而成本感知的 GRPO 则探索协调策略,以便在长视野交互期间选择性读取、更新和整合该状态。在搭载 Qwen3-8B LLM 的 ALFWorld 上实例化后,EvoHarness-RL 达到 96.9% 的成功率,并展现出两种关键动态:管控器退火,即训练将重复的管控器使用模式内化到模型策略中,使智能体从频繁调用管控器转向选择性访问外部状态;管控器进化,即进度更新和经验整合将管控器优化为紧凑、任务自适应的状态基底。这些结果表明,长视野智能体可从用于构建和协调外部管控器工作空间的可训练策略中受益,而非仅依赖更强的工具或更大的内存。
英文摘要
Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, recover from failures, and reuse experience across extended interactions. Yet existing harnesses and their use are often tailored to environments and controlled through prompts, heuristics, or system-specific rules, making agent and harness coordination difficult to jointly optimize. We introduce EvoHarness-RL, a unified framework that separates environment-specific harness implementations from a shared policy-facing interface. EvoHarness-RL organizes external support into a Belief, Progress, and Experience (BPE) workspace and exposes four compact harness actions for accessing and updating this state. We first instantiate BPE as an inference-time scaffold and then make harness coordination learnable through supervised initialization followed by cost-aware GRPO. Across heterogeneous long-horizon tasks, EvoHarness-Base improves the average success rate of frontier models by 10.0 percentage points, while EvoHarness-RL outperforms the strongest open-source baseline by 8.5 percentage points, together with higher RL rollout efficiency and stronger generalization to unseen tasks. Our analyses show that training gradually shifts agents from frequent scaffold use toward selective, environment-dependent harness access as the policy becomes more capable, while the external workspace continues to evolve and refine itself to better support task execution and generalization. Together, these results show that long-horizon agents benefit not only from external scaffolding itself, but also from learning how and when to coordinate with external support as part of a cost-aware policy.
CommentsAccepted to LLA@COLM 2026