arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LoGRA:利用低秩梯度草图扩展大语言模型强化学习

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

Shaokun Zhang, Yifan Zhang, Jian Hu, Yueying Li, Hao Zhang, Binfeng Xu, Jan Kautz, Yi Dong

arXiv 2610.06647首次发表:更新:

AI 中文总结

LoGRA通过低秩梯度草图压缩内存并配合预测KL步长控制,在推理任务中平均减少45.7%训练内存,实现270亿参数模型稳定训练超1100步。

AI 中文摘要

强化学习(RL)极大地提升了大语言模型(LLM)的能力,但其内存需求仍是更广泛采用的障碍。我们提出LoGRA,一种用于RL后训练的方法,通过将有用的学习信号保留在低秩梯度草图中来减少内存占用。这些紧凑表示同时支持模型更新和高效策略同步。为防止过大的更新干扰学习,我们将梯度压缩与预测KL步长控制相结合,该控制在每个更新应用前估计策略变化并相应调整其幅度。在推理任务中,LoGRA平均减少训练内存高达45.7%,且不牺牲性能。它还使得在单个八GPU节点上稳定训练270亿参数模型超过1100步成为可能,而稠密Adam在此场景下会内存耗尽,从而使此前因内存不足而不可行的RL训练变得切实可行。代码可在Molt库中获取。

英文摘要

Reinforcement learning has greatly advanced the capabilities of large language models, but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. With all techniques combined, LoGRA reduces average training memory usage by up to 45.7% across reasoning tasks without compromising performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra}{Molt library}.

Comments16 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑