发表机构
University of Miami; Harvard University(迈阿密大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出有界陈旧性异步进化策略,在长时程智能体任务上验证其与同步ES性能相当,且能容忍策略滞后,为ES后训练异步化奠定基础。
AI 中文摘要
基于LLM的长时程智能体后训练常常受制于轨迹生成:轨迹跨越多个交互轮次,完成时间差异显著,同步更新屏障使较快的工作者等待掉队者。异步强化学习已在LLM后训练中采用,通过按到达顺序消费轨迹来解决这一低效问题,但引入了策略滞后和离策略优化。进化策略(ES)为LLM后训练提供了一种免反向传播的替代方案,然而它依赖更多数量的轨迹,且现有实践大多保持同步。在这篇短文形式论文中,我们引入了有界陈旧性的异步进化策略,并在Endless Terminals基准上使用Qwen2.5-7B-Instruct进行演示。在三个评估种子下,自然Async-1与同步ES表现相当,分别达到25.9%与25.4%的保留集成功率。控制进度表将每个更新批次中的10%延迟四个或八个策略更新,分别仅使成功率降低1.6和3.1个百分点,而无需显式离策略校正。GRPO整体表现更佳,达到29.0%的保留集成功率,但重要的是,我们的结果表明ES能容忍适度的策略陈旧性且退化有限,为未来基于异步算法的ES后训练改进开辟了可能性。据我们所知,我们是首个在多轮终端风格智能体编码任务上展示同步和异步ES有效性的工作。
英文摘要
LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-training addresses this inefficiency by consuming trajectories as they arrive, but introduces policy lag and off-policy optimization. Evolution strategies (ES) offer a backpropagation-free alternative for LLM post-training, yet it relies on a larger number of rollouts and existing practices have remained largely synchronous. In this short-form paper, we introduce bounded-staleness asynchronous ES and demonstrate it on Endless Terminals benchmark using Qwen2.5-7B-Instruct. Across three evaluation seeds, natural Async-1 matches synchronous ES, achieving 25.9\% versus 25.4\% held-out success. Controlled schedules that delay 10\% of each update cohort by four or eight policy updates reduce success by only 1.6 and 3.1 percentage points, respectively, without explicit off-policy correction. GRPO performs better overall, reaching 29.0\% held-out success, but importantly our results show that ES tolerates moderate policy staleness with limited degradation, opening possibilities for future improvement of ES-based post-training with asynchronous algorithms. To the best of our knowledge, we are the first to demonstrate the effectiveness of sync and async ES on a multi-turn terminal style agentic coding task.