AI 中文总结
该研究推出开源库rl-triton,用Triton实现统一关联扫描框架,将7种RL估计算法统一为一阶线性递推,在大规模并行模拟中实现1.6至5.7倍的速度提升。
AI 中文摘要
我们推出rl-triton,这是一个用Triton实现的、用于强化学习信用分配的高性能GPU内核开源库。核心贡献是一个统一的关联扫描框架,它将7种不同的RL估计算法——广义优势估计(GAE)、V-Trace、Retrace(λ)、TD(λ)回报、折扣回报、资格迹以及情节前缀和——重新表述为单个一阶线性递推的实例,该递推可在O(log T)的并行步骤中求解。所有算法共享相同的关联扫描算子,特定算法的融合Triton内核会在片上构建它们的递推系数。我们通过代数方法验证了该关联算子,并明确了终止和截断情节的处理方式。基准测试显示,在大规模并行模拟场景(数千个环境、短回合)中,与向量化baseline相比,全调用速度提升了1.6至5.70倍。该范围涵盖了7种算法在两种GPU上的表现,以及是否处理每步截断的情况。对于大多数算法,随着序列长度增加,速度提升会增大,因为随着log T增长,baseline需要更多扫描阶段,每个阶段都会增加一次中间高带宽内存(HBM)往返。该库可在指定链接获取。
英文摘要
We present rl-triton, an open-source library of high-performance GPU kernels for reinforcement learning credit assignment, implemented in Triton. The core contribution is a unified associative scan framework that recasts seven distinct RL estimation algorithms - Generalized Advantage Estimation (GAE), V-Trace, Retrace($λ$), TD($λ$) returns, discounted returns, eligibility traces, and episodic prefix sums - as instances of a single first-order linear recurrence solved in $O(\log T)$ parallel steps. All algorithms share the same associative scan operator, with algorithm-specific fused Triton kernels constructing their recurrence coefficients on-chip. We verify the associative operator algebraically and define the treatment of terminated and truncated episodes explicitly. Benchmarks show a 1.6-5.70$\times$ full-call speedup over a vectorized torch-compile baseline in the massively parallel simulation regime (thousands of environments, short rollouts). The reported range covers all seven algorithms on both GPUs, both with and without per-step truncation handling. For most algorithms, speedups increase at longer sequence lengths, as the baseline requires more scan stages as $\log T$ grows, each adding an intermediate HBM round-trip. The library is available at https://github.com/simonsays1980/rl-triton.
Comments18 pages, 3 figures, 6 tables. Code: https://github.com/simonsays1980/rl-triton