HISPO:基于熵导出分段的层次重要性采样策略优化
HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
- Cho Chun Shik Graduate School of Mobility(赵春植移动研究生院)
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
HISPO提出分段级重要性采样策略优化,利用熵导出分段和软显著性权重,在长格式数学推理中优于GRPO和GSPO,提升RLVR性能。
AI中文摘要:
具有可验证奖励的强化学习(RLVR)已成为提升语言模型数学推理能力的核心方法,但长格式补全引入了困难的信用分配问题:解决方案轨迹的不同部分可能对最终正确性的贡献不均。现有的用于RLVR的策略优化目标通常在词元级别(GRPO、DAPO)或序列级别(GSPO)应用重要性采样校正,从而在响应中施加了不同粒度的信用分配。我们提出了层次重要性采样策略优化(HISPO),一种分段级策略优化方法,该方法构建基于熵的连续分段(在 rollout 时生成),分配基于熵的软显著性权重,并在分段粒度上应用裁剪的重要性采样校正。这提供了介于词元级 GRPO/DAPO 与序列级 GSPO 之间的中间校正单元。我们通过在数学推理任务上微调 Qwen3-1.7B-Base 来评估 HISPO。在六个基准测试中,HISPO 在所有基准上相对于最强基线提升了 Pass@8,并在其中五个基准上达到或超过了最强基线的 Acc@8。在 AIME25 上,HISPO 相对于 GRPO 提升了 +3.75 Acc@8 和 +3.78 Pass@8,相对于 GSPO 提升了 +2.50 Acc@8 和 +1.27 Pass@8。这些结果表明,分段级校正是长格式数学推理中 RLVR 的一种有前景的粒度。
英文摘要:
Reinforcement learning with verifiable rewards (RLVR) has become a central approach for improving mathematical reasoning in language models, but long-form completions introduce a difficult credit-assignment problem: different parts of a solution trace may contribute unevenly to final correctness. Existing policyoptimization objectives for RLVR commonly apply importance-sampling correction at either the token level (GRPO, DAPO) or the sequence level (GSPO), imposing different granularities for assigning credit across a response. We introduce Hierarchical Importance-Sampling Policy Optimization (HISPO), a segment-level policy-optimization method that constructs rollout-time entropy-derived contiguous segments, assigns soft entropy-based saliency weights, and applies clipped importance-sampling correction at the segment granularity. This provides an intermediate correction unit between token-level GRPO/DAPO and sequence-level GSPO. We evaluate HISPO by fine-tuning Qwen3-1.7B-Base on mathematical reasoning tasks. Across six benchmarks, HISPO improves Pass@8 over the strongest baseline on all benchmarks and matches or exceeds the strongest baseline in Acc@8 on five of them. On AIME25, HISPO improves over GRPO by +3.75 Acc@8 and +3.78 Pass@8, and over GSPO by +2.50 Acc@8 and +1.27 Pass@8. These results suggest that segment-level correction is a promising granularity for RLVR in long-form mathematical reasoning.