发表机构
University of Southern Queensland(南昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HiLoRe通过策略更新暴露度分配存储、压缩与重计算,在GRPO训练中实现高达13.5%的吞吐提升,同时保持下游性能。
AI 中文摘要
组相对策略优化(GRPO)使得学习者侧激活成为主要的内存-计算瓶颈:梯度检查点通过重计算减少激活内存,但固定调度在48-GB GPU上可能留下约18 GB未使用,尽管存在大量重计算开销。现有的激活管理方法根据执行成本、张量属性或通用压缩敏感性来设定状态保真度,未明确将GRPO的解析更新结构纳入状态保真度分配。我们将这种依赖形式化为策略更新暴露度,将当前GRPO损失系数与状态级近似敏感性联系起来。这些系数在反向传播之前无需额外反向传播即可获得。我们提出HiLoRe,它利用测量的恢复效用和更新条件近似风险,在图属性恢复单元中分配高精度存储、低精度压缩和确定性重计算。它在校准的风险预算下,将高精度存储和确定性重计算与低精度恢复相结合。在五个模型-任务设置中,具有2K响应且内存小于GC每GPU演员更新峰值的1.10倍,HiLoRe的演员更新吞吐量提升相对于GC达到13.5%,相对于评估的最快基线达到7.9%,配对平均下游分数差异低于0.6个百分点。
英文摘要
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.
CommentsUnder Review