arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

All Work And No Play Makes Jack a Dull Boy: 理解并防止RLVR中的灾难性策略崩溃

All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR

Qiyuan Huang, Tianshi Xu, Meng Li

arXiv 2610.02835首次发表:更新:

发表机构

Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示RLVR中GRPO算法导致策略容量收缩引发灾难性崩溃,提出网格学习暴露多策略并防主导,在多个基准上提升达13.4个百分点。

AI 中文摘要

在大型语言模型(LLMs)使用可验证奖励强化学习(RLVR)进行后训练期间,GRPO风格的算法可能表现出严重的后期崩溃。基于提示的探测显示,这并非良性的策略修剪,而是有效策略能力的有害收缩,使得不同的推理策略越来越难以访问。为了表征这一现象,我们通过轨迹级别的策略更新交互来定义策略,并开发了一个结合优化动力学和信息论的统一理论框架。我们证明了主要的RLVR目标会逐渐将概率质量集中到单一策略上,而维持非平凡的任务准确率需要最小的策略容量。这两个结果之间的冲突为灾难性崩溃提供了机制性解释。我们进一步推导出镜像纠缠指数(MEI)作为轻量级的在线预警信号。为了防止崩溃,我们提出了网格学习(Mesh Learning),它暴露多种推理策略并防止任何单一策略主导优化。在AIME26、AIME25、MATH-500、GPQA和LiveCodeBench上,Mesh Learning在Qwen和Phi模型系列中始终优于强基线,分别提高了高达13.4个百分点和11.5个百分点。这些结果确立了策略保留作为稳定RLVR的关键原则。代码可在该https URL获取。

英文摘要

During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.

Comments84 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑