arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PPO-HSC:基于广域策略覆盖优化的探索性强化学习框架

PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization

Yujie Shen, Haowen Chen

arXiv 2607.16206首次发表:更新:

发表机构

School of Mathematics, Hunan University; College of Computer Science and Electronic Engineering, Hunan University(湖南大学数学学院; 湖南大学计算机科学与电子工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PPO-HSC是用于解决大语言模型微调中模式崩溃问题的探索性强化学习框架,通过纳入高阶采样覆盖奖励及维护动态轨迹库,提高解决方案多样性与状态空间覆盖,在数学推理和代码生成任务中表现出色。

AI 中文摘要

本文介绍了PPO-HSC(近端策略优化与高阶采样覆盖),这是一个探索性强化学习框架,旨在解决大语言模型微调中模式崩溃的“无形枷锁”问题。标准的基于可验证奖励的强化学习(RLVR)虽能强化高奖励轨迹,但常使模型过度优化已知解决方案。PPO-HSC纳入高阶采样覆盖(HSC)奖励,激励发现“低相似性但高效性”的推理模式。通过维护经过验证的独特解决方案的动态轨迹库,该框架提供了一个可微信号,奖励语义新颖性并通过合理性约束确保结构合理性。对数学推理(GSM8K、SVAMP)和代码生成任务的实证评估表明,PPO-HSC显著提高了解决方案多样性和状态空间覆盖,同时保持或超越了现有强化学习基线的准确性和语法完整性。

英文摘要

This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning framework designed to address the "Invisible Shackles" of mode collapse in Large Language Model (LLM) fine-tuning. While standard Reinforcement Learning from Verifiable Rewards (RLVR) effectively reinforces high-reward trajectories, it often leads models to over-optimize known solutions, sacrificing curiosity and the ability to explore broader solution manifolds. To overcome this, PPO-HSC incorporates a High-order Sampling Coverage (HSC) reward that incentivizes the discovery of "low-similarity yet high-validity" reasoning patterns. By maintaining a dynamic trajectory library of verified unique solutions, the framework provides a differentiable signal that rewards semantic novelty while ensuring structural rationality through a plausibility constraint. Empirical evaluations on mathematical reasoning (GSM8K, SVAMP) and code generation tasks demonstrate that PPO-HSC significantly enhances solution diversity and state-space coverage while maintaining or surpassing the accuracy and syntax integrity of state-of-the-art RL baselines.

Comments13 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑