发表机构
Pohang University of Science and Technology (POSTECH)(浦项科技大学(POSTECH))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RLVR中推理覆盖不足的问题,提出难度自适应句子熵引导树结构策略优化(DATPO),通过树搜索与多样性优势项扩展pass@k,在数学推理基准上取得更优性能。
AI 中文摘要
基于可验证奖励的强化学习(RLVR)是近期大型推理模型成功的关键。然而,尽管RLVR显著提高了单样本准确率,但由于训练期间探索有限,它往往无法扩展模型的内在推理覆盖范围(pass@k)。为解决这一问题,我们优化了训练时rollout的结构设计以提升pass@k。我们的分析确定了三个关键设计原则:(1)难度自适应rollout在扩展pass@k方面可发挥重要作用,而不仅仅是作为效率启发式;(2)基于树的rollout在发现正确答案方面优于并行采样;(3)句子熵引导的分叉克服了token级分支的局部化现象,以最大化语义多样性。基于这些见解,我们提出了DATPO(难度自适应句子熵引导树结构策略优化)。DATPO将难度自适应树搜索与兄弟多样性优势项相结合,明确促进语义多样性以在训练期间扩展推理覆盖范围。在数学推理基准上的实验表明,DATPO在pass@k方面尤其优于基线,这直接转化为更优的测试时扩展性能。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has been central to the recent success of Large Reasoning Models. However, while RLVR significantly improves single-sample accuracy, it often fails to expand the model's intrinsic reasoning coverage (pass@k) due to limited exploration during training. To address this, we optimize the structural design of train-time rollouts to enhance pass@k. Our analysis identifies three key design principles: (1) difficulty-adaptive rollout can play an important role in expanding pass@k, beyond serving as an efficiency heuristic; (2) tree-based rollout outperforms parallel sampling in discovering correct answers; and (3) sentence-entropy-guided forking overcomes the localization phenomenon of token-level branching to maximize semantic diversity. Building on these insights, we propose DATPO (Difficulty-Adaptive Sentence-entropy-guided Tree-structured Policy Optimization). DATPO integrates difficulty-adaptive tree search with a sibling-diversity advantage term, explicitly promoting semantic diversity to expand reasoning coverage during training. Experiments on mathematical reasoning benchmarks demonstrate that DATPO outperforms baselines especially in pass@k, which directly translates to superior test-time scaling performance.