arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GCPO:针对大语言模型的rollout强化学习中子空间几何的诊断与约束

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao

arXiv 2608.11674首次发表:更新:

发表机构

Shanghai AI Lab(上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GCPO通过几何约束策略优化方法,诊断rollout强化学习的子空间几何问题,在大语言模型后训练中提升了多任务性能并解决了训练不稳定等问题

AI 中文摘要

GRPO等on-policy rollout方法是大语言模型后训练的核心,但它们常存在训练不稳定、跨任务能力下降、响应长度膨胀的问题。尽管已有工作对聚合更新的子空间几何进行了表征,但该几何的逐步变化及其与模型性能的关系仍不清楚。我们提出Principal-Subspace Overlap,一种针对单个rollout更新相对于预训练权重主导奇异子空间的维度校正度量。尽管平均重叠度较低,但性能下降前常出现瞬时峰值。为解决该问题,我们提出GCPO(Geometrically Constrained Policy Optimization,几何约束策略优化),其应用硬双侧正交投影将更新约束到互补子空间,从结构上防止此类偏移。在Qwen3-8B和GLM4-9B上的数学推理、代码生成及工具使用任务中,GCPO的性能始终优于GRPO及DAPO、GSPO等近期变体,较基础模型和最强基准分别提升最高27.69和2.37个百分点。此外,GCPO可保留通用能力、消除响应长度膨胀并稳定策略熵。我们的发现为稳定强化学习后训练提供了新的诊断视角和原则性设计思路。

英文摘要

On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We introduce Principal-Subspace Overlap, a dimension-corrected measure of individual rollout updates relative to the dominant singular subspaces of pretrained weights. Despite low average overlap, transient spikes often precede performance degradation. To address this, we propose GCPO (Geometrically Constrained Policy Optimization), which applies hard bilateral orthogonal projections to constrain updates to the complementary subspaces, preventing such excursions by construction. Across mathematical reasoning, code generation, and tool-use tasks on Qwen3-8B and GLM4-9B, GCPO consistently outperforms GRPO and recent variants, including DAPO and GSPO, improving over the base models and the strongest baseline by up to 27.69 and 2.37 points, respectively. Furthermore, GCPO preserves general capabilities, eliminates response-length inflation, and stabilizes policy entropy. Our findings provide a new diagnostic lens and a principled design perspective for stable reinforcement learning post-training.

Comments15 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑