AI 中文总结
psRL是利用训练样本前缀冗余的智能体AI训练系统,通过新颖前缀共享机制与KV缓存管理器提升分布式训练效率,吞吐量较现有系统最高提升5.2倍。
AI 中文摘要
在现代智能体AI训练中,系统瓶颈正从rollout(试跑)阶段转向update(更新)阶段。树状和分步强化学习(RL)等新兴采样策略大幅增加了训练样本量,同时边际rollout成本相对较低,导致更新阶段主导了端到端执行时间。关键的是,这一转变暴露了新的优化机会,因为生产轨迹显示训练样本间存在大量前缀冗余。本文提出psRL(即前缀共享强化学习),这是一种利用训练样本间前缀冗余的智能体AI训练系统。借助更新阶段固有的全局可见性和数据不变性,psRL实现了分布式训练的高效工作负载调度与内存管理。具体而言,psRL引入两种新颖的前缀共享机制,可在GPU工作节点间实现灵活、细粒度的工作负载分配,同时优化前缀复用并实现负载均衡。此外,psRL实现了一种新的底层KV缓存管理器,支持自适应块大小分配与动态KV缓存,在维持高前缀命中率的同时最大化内存利用率。基于生产轨迹的评估表明,psRL的吞吐量比现有系统高出最多5.2倍,其源代码即将公开。
英文摘要
In modern agentic AI training, the system bottleneck is shifting from rollout to update. Emerging sampling strategies such as tree-structured and step-wise RL greatly increase training sample volume while incurring relatively low marginal rollout cost, causing the update phase to dominate the end-to-end execution time. Crucially, this shift exposes a new optimization opportunity, as production traces reveal substantial prefix redundancy across training samples. In this paper, we propose psRL (prefix sharing for RL), a new training system for agentic AI designed to exploit prefix redundancy among training samples. Leveraging the global visibility and data immutability inherent to the update phase, psRL achieves efficient workload scheduling and memory management for distributed training. Specifically, psRL introduces two novel prefix-sharing mechanisms that enable flexible, fine-grained workload distribution across GPU workers, simultaneously optimizing prefix reuse and achieving load balancing. Moreover, psRL implements a new underlying KV cache manager that facilitates adaptable block-size allocation and dynamic KV caching, maximizing memory utilization while maintaining a high prefix hit rate. Evaluations using production traces demonstrate that psRL outperforms existing systems by up to 5.2x in throughput. The source code will be publicly available soon.
Comments15 pages, 15 figures, 2 table