arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27420cs.CL

通过RLVR中的弱模型引导增强大语言模型探索能力

Boosting LLM Exploration via Weak-Model Guidance in RLVR

Xingyu Shen, Huishuai Zhang, Peng Li, Yinchun Wang, Dongyan Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出弱模型引导的RLVR方法,通过弱模型的部分推理轨迹增强LLM探索,缓解熵崩溃,在数学基准上提升大k值下的推理性能与覆盖范围。

中文摘要 AI 辅助

带可验证奖励的强化学习(RLVR)显著提升大语言模型(LLM)的推理能力,但常导致策略熵下降,进而缩小推理覆盖范围并降低大k值下的pass@k指标。现有方法虽通过算法正则化缓解这种熵崩溃,但忽略了跨模型非参数扰动的作用。本研究提出一种简单却有效的方法,用于在RLVR过程中保留LLM的生成多样性:不再仅依赖内部探索,而是强制目标模型基于较小规模的弱语言模型生成的部分推理轨迹生成答案。这些陌生的前缀能有效打破过自信状态,鼓励探索不同的推理路径。我们通过实证研究外部前缀的潜力,揭示了分布差异对RLVR训练中探索动态的影响机制。在多个数学基准上的实验表明,本方法始终优于普通RLVR;值得注意的是,随着k值增大,性能提升愈发显著,证明推理覆盖范围得到大幅扩展。此外,本方法能高效缓解熵崩溃,无需额外的监督微调(SFT)、复杂的奖励设计或繁琐的提示工程。

英文摘要

Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet effective approach to preserve the generative diversity of LLMs during RLVR. Instead of relying solely on internal exploration, we force the target model to generate answers based on partial reasoning trajectories generated by a smaller, weaker language models. These unfamiliar prefixes effectively disrupt over-confidence and encourage the exploration of distinct reasoning paths. We empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training. Experiments across multiple mathematical benchmarks show that our method consistently outperforms vanilla RLVR. Notably, the performance gain becomes increasingly pronounced as $k$ scales up, demonstrating a substantial expansion of reasoning coverage. Furthermore, our approach efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

发表机构

  • Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
  • National Engineering Research Center of New Electronic Publishing Technologies(国家新型电子出版技术工程研究中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑