arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17299cs.LGcs.AIcs.CLcs.OS

WAR:用于同步智能体强化学习的工作负载感知展开

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning

Ryan Xu, Atlas Zhao, David Bao, Frank Du

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对智能体强化学习中展开成为系统瓶颈的问题,提出工作负载感知展开系统WAR,通过联合优化解码和调度加速同步智能体RL,在不同负载下提高吞吐量,消除主要展开瓶颈,为长上下文智能体训练提供实用路径。

中文摘要 AI 辅助

在智能体强化学习(RL)中,长期展开生成已成为主要的系统瓶颈。随着智能体与环境多轮交互,轨迹迅速增长到数万个令牌,使得同步RL训练越来越受展开的限制。我们提出了WAR,一种工作负载感知展开系统,通过联合优化解码和调度,大幅加速同步智能体RL。WAR基于一个关键观察:最优展开优化策略取决于运行时负载。在低负载下,WAR使用后缀解码实现无模型推测性解码,在高负载下,WAR将优化重点转移到缓存感知调度。通过结合解码级后缀重用和系统级展开调度,WAR在不改变底层RL算法的情况下,显著提高了吞吐量。在低负载下,WAR将长上下文智能体展开吞吐量提高了1.4倍,在高负载下提高了1.6倍。这些结果表明,WAR消除了同步智能体RL中的一个主要展开瓶颈,并为可扩展的长上下文智能体训练提供了一条实用路径。

英文摘要

Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with environments over many turns, trajectories rapidly grow to tens of thousands of tokens, making synchronous RL training increasingly constrained by rollout. We propose WAR, a workload-aware rollout system that substantially accelerates synchronous agentic RL by jointly optimizing decoding and scheduling. WAR is built on a key observation: the optimal rollout optimization strategy depends on runtime load: (1) Under low load, WAR enables model-free speculative decoding with SuffixDecoding, which reuses suffix patterns from previously completed trajectories as speculative drafts for future rollouts. Unlike model-based drafters, SuffixDecoding introduces no additional draft model and avoids GPU contention with rollout generation. (2) Under high load, where saturated batched decoding leaves limited room for speculative speedup, WAR shifts the optimization focus to cache-aware scheduling. A global scheduler places requests across rollout replicas based on cache locality, trajectory progress and server load, reducing redundant KV-cache recomputation and mitigating load imbalance. By combining decoding-level suffix reuse with system-level rollout scheduling, WAR delivers robust throughput improvements across workload regimes without changing the underlying RL algorithm. WAR improves long-context agentic rollout throughput by 1.4x under low load and up to 1.6x under high load. These results show that WAR removes a major rollout bottleneck in synchronous agentic RL and provides a practical path toward scalable long-context agent training.

↑