发表机构
School of Mathematics, Hunan University; College of Computer Science and Electronic Engineering, Hunan University(湖南大学数学学院; 湖南大学计算机科学与电子工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PPO-HSC是用于解决大语言模型微调中模式崩溃问题的探索性强化学习框架,通过纳入高阶采样覆盖奖励及维护动态轨迹库,提高解决方案多样性与状态空间覆盖,在数学推理和代码生成任务中表现出色。
AI 中文摘要
本文介绍了PPO-HSC(近端策略优化与高阶采样覆盖),这是一个探索性强化学习框架,旨在解决大语言模型微调中模式崩溃的“无形枷锁”问题。标准的基于可验证奖励的强化学习(RLVR)虽能强化高奖励轨迹,但常使模型过度优化已知解决方案。PPO-HSC纳入高阶采样覆盖(HSC)奖励,激励发现“低相似性但高效性”的推理模式。通过维护经过验证的独特解决方案的动态轨迹库,该框架提供了一个可微信号,奖励语义新颖性并通过合理性约束确保结构合理性。对数学推理(GSM8K、SVAMP)和代码生成任务的实证评估表明,PPO-HSC显著提高了解决方案多样性和状态空间覆盖,同时保持或超越了现有强化学习基线的准确性和语法完整性。
英文摘要
This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the "Invisible Shackles" of mode collapse in Large Language Model (LLM) fine-tuning. While standard Reinforcement Learning from Verifiable Rewards (RLVR) effectively reinforces high-reward trajectories, it often leads models to over-optimize known solutions, sacrificing curiosity and the ability to explore broader solution manifolds. To overcome this, PPO-HSC incorporates a High-order Sampling Coverage (HSC) reward that incentivizes the discovery of "low-similarity yet high-validity" reasoning patterns. By maintaining a dynamic trajectory library of verified unique solutions, the framework provides a differentiable signal that rewards semantic novelty while ensuring structural rationality through a plausibility constraint. Empirical evaluations on mathematical reasoning (GSM8K, SVAMP) and code generation tasks demonstrate that PPO-HSC significantly enhances solution diversity and state-space coverage while maintaining or surpassing the accuracy and syntax integrity of state-of-the-art RL baselines.
Comments13 pages, 3 figures