发表机构
Tsinghua University; Zhongguancun Laboratory(清华大学; 中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TideRL是一种就绪感知弹性RL系统,通过CTB、RA²P和ERS技术,在多轮智能体工作负载上大幅提升RL训练有效吞吐量,同时优化KV缓存命中率、训练时间和等待时间。
AI 中文摘要
针对大语言模型的强化学习(RL)正朝着多轮智能体工作负载发展,其中回滚任务会反复暂停以等待外部环境,随上下文增长而恢复,并在高度可变的时间完成。在这种场景下,以训练吞吐量衡量的RL训练有效吞吐量(goodput)比原始GPU利用率更重要:GPU等待和重复的预填充重新计算属于纯开销。我们提出TideRL,这是一种具备连续任务批处理(CTB)、资源感知参考-演员流水线(RA²P)和弹性资源缩放(ERS)的就绪感知弹性RL系统。CTB保留有用的回滚状态,RA²P从就绪积压和到达间隔中选择解耦流或共位置聚合,ERS利用相同的就绪信号在回滚和训练之间移动秩。在纯文本和多模态智能体工作负载上,TideRL与同步基线相比将RL训练有效吞吐量提升高达5.6倍,与异步基线相比提升超过33%,同时达到相似的任务性能。它还将KV缓存命中率提升1.58倍,将每步训练时间减少高达44.3%,并将总等待时间减少高达77.6%。
英文摘要
Reinforcement learning (RL) for large language models is moving toward multi-turn agentic workloads, where rollout tasks repeatedly pause for external environments, resume with growing contexts, and finish at highly variable times. In this setting, RL training goodput, measured by training throughput, matters more than raw GPU occupancy: GPU waiting and repeated prefill recomputation are pure overhead. We present TideRL, a readiness-aware elastic RL system with Continuous Task Batching, Resource-Aware Ref-Actor Pipelining, and Elastic Resource Scaling. CTB preserves useful rollout state, $\textrm{RA}^2\textrm{P}$ selects between decoupled streaming and colocated aggregation from the ready backlog and arrival interval, and ERS moves ranks between rollout and training using the same readiness signals. Across text-only and multi-modal agentic workloads, TideRL improves RL training goodput by up to 5.6$\times$ over synchronous baselines and over 33% over asynchronous baselines, while reaching similar task performance. It also improves KV cache hit rate by 1.58$\times$, reduces per-step training time by up to 44.3%, and cuts total waiting time by up to 77.6%.