arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HiLoRe:高效GRPO训练中存储、压缩或重计算的选择

HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training

Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang

arXiv 2609.33570首次发表:更新:

发表机构

University of Southern Queensland(南昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HiLoRe通过策略更新暴露度分配存储、压缩与重计算,在GRPO训练中实现高达13.5%的吞吐提升,同时保持下游性能。

AI 中文摘要

组相对策略优化(GRPO)使得学习者侧激活成为主要的内存-计算瓶颈:梯度检查点通过重计算减少激活内存,但固定调度在48-GB GPU上可能留下约18 GB未使用,尽管存在大量重计算开销。现有的激活管理方法根据执行成本、张量属性或通用压缩敏感性来设定状态保真度,未明确将GRPO的解析更新结构纳入状态保真度分配。我们将这种依赖形式化为策略更新暴露度,将当前GRPO损失系数与状态级近似敏感性联系起来。这些系数在反向传播之前无需额外反向传播即可获得。我们提出HiLoRe,它利用测量的恢复效用和更新条件近似风险,在图属性恢复单元中分配高精度存储、低精度压缩和确定性重计算。它在校准的风险预算下,将高精度存储和确定性重计算与低精度恢复相结合。在五个模型-任务设置中,具有2K响应且内存小于GC每GPU演员更新峰值的1.10倍,HiLoRe的演员更新吞吐量提升相对于GC达到13.5%,相对于评估的最快基线达到7.9%,配对平均下游分数差异低于0.6个百分点。

英文摘要

Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory < 1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑