超越欧几里得裁剪:通过黎曼等距策略优化克服大语言模型强化学习中的探索崩溃
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
浏览论文内容
中文总结 AI 辅助
研究大语言模型强化学习中PPO-Clip的探索崩溃问题,提出黎曼等距策略优化(RIPO)方法,通过保证黎曼流形上等距策略更新平衡探索与利用,在七个竞争级基准上显著超越现有算法。
中文摘要 AI 辅助
强化学习(RL)已成为增强大语言模型推理能力的主导范式。然而,基于近端策略优化裁剪(PPO-Clip)的RL算法存在探索崩溃问题。后续工作主要是启发式的,未找到PPO-Clip失败的根本原因。本文揭示了PPO-Clip的根本缺陷:它使用欧几里得度量隐式测量策略差异,与策略黎曼流形上的内在几何不一致,导致探索崩溃。为此提出黎曼等距策略优化(RIPO),保证在黎曼流形上进行等距策略更新,平衡探索与利用。实验表明RIPO在七个竞争级基准上显著超越现有算法。
英文摘要
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fundamental flaw of PPO-Clip: it implicitly measures policy discrepancy using Euclidean metric, which is theoretically inconsistent with the intrinsic geometry on the policy Riemannian manifold. This geometric mismatch results in overly conservative updates in low-probability regions while aggressive in high-probability regions, ultimately collapsing exploration. To correct this geometric flaw, we propose Riemannian Isometric Policy Optimization (RIPO), which guarantees isometric policy updates on the Riemannian manifold, effectively balancing exploration and exploitation. We further show that RIPO achieves a favorable bias-variance trade-off, which stabilizes optimization. Extensive experiments demonstrate that RIPO significantly surpasses existing LLM RL algorithms across seven competition-level benchmarks (up to 60% improvement over GRPO on AIME24).