ThinkPrior:RLVR中冷启动提示选择的零回滚难度先验
ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
浏览论文内容
中文总结 AI 辅助
ThinkPrior利用外部锚点构建零回滚难度先验,在RLVR中冷启动选择提示,减少早期静默组和浪费回滚,且不损害最终准确率。
中文摘要 AI 辅助
在使用组相对策略优化(GRPO)训练的可验证奖励强化学习(RLVR)中,本文研究的无KL奖励优势项依赖于组内奖励的变异性。如果一组中的所有回滚都正确或全部错误,则其组相对优势完全相同为零;这些零优势的静默组不提供任何奖励优势梯度,然而均匀采样将一次运行中39%的回滚花费在这些组上。基于历史的提示选择必须首先花费目标策略回滚来估计难度,从而产生带有回滚浪费的冷启动;ThinkPrior则相反,在第一次目标策略回滚之前,通过一次离线过程使用外部锚点来构建零回滚难度先验。验证器评分的锚点通过率提供了Beta后验的外部锚点初始化;ThinkPrior根据预期可学习性进行选择,然后根据训练结果进行更新,既不改变损失函数也不改变优化器。在Qwen2.5-Math-7B上跨十六个随机种子,ThinkPrior将早期静默组减少了一半以上,并将到第30步的浪费回滚削减了近五分之一,同时我们未检测到最终准确率的差异。在这个250提示池中,固定预算的结果是重新分配而非净节省。所测得的ThinkPrior+DAPO组合将生成的回滚减少了10.6%,而两组都保持相同的3840回滚更新预算。该先验在首次选择之前不需要目标策略回滚,但之后的后续先验使用目标策略的结果。
英文摘要
In reinforcement learning with verifiable rewards (RLVR) trained with group relative policy optimization (GRPO), the KL-free reward-advantage term studied here depends on within-group reward variation. If all rollouts in a group are correct or all are wrong, their group-relative advantages are identically zero; these zero-advantage silent groups provide no reward-advantage gradient, yet uniform sampling spends 39% of a run's rollouts on them. History-based prompt selection must first spend target-policy rollouts to estimate difficulty, creating a cold start with rollout waste; ThinkPrior instead uses an external anchor in one offline pass to construct a zero-rollout difficulty prior before the first target-policy rollout. The verifier-scored anchor pass rate supplies an external-anchor initialization for a Beta posterior; ThinkPrior selects by expected learnability and then updates from training outcomes, changing neither the loss nor the optimizer. On Qwen2.5-Math-7B across sixteen seeds, ThinkPrior more than halves early silent groups and cuts wasted rollouts through step 30 by nearly a fifth, while we detect no difference in final accuracy. On this 250-prompt pool the fixed-budget result is a reallocation rather than a net saving. The measured ThinkPrior+DAPO composition reduces generated rollouts by 10.6% while both arms retain the same 3840-rollout update budget. The prior requires no target-policy rollout before the first selection, but the posterior thereafter uses target-policy outcomes.
发表机构
- Stony Brook University(石溪大学)
- University of Minnesota Twin Cities(明尼苏达大学双城校区)
机构由 AI 辅助整理,请以论文原文为准。