AI 中文总结
LEEPS是一种潜在引导探索-利用提示采样器,通过自适应分配rollout预算等方式优化提示采样,在6个数学推理基准和3个OOD通用推理基准上均取得最优性能,且训练开销较小。
AI 中文摘要
带可验证奖励的强化学习(RLVR)可提升大语言模型的推理能力,但具有相同rollout奖励的提示组会消耗生成预算却无法提供有效学习信号。预rollout提示选择可通过在rollout生成前筛选提示减少这种浪费,不过现有预rollout方法难以平衡探索与利用:反复利用历史上有信息的提示会缩小训练覆盖范围,而更广泛的探索则会降低有信息提示的占比。为解决这些局限,我们提出LEEPS,即潜在引导探索-利用提示采样器,它能自适应平衡对先前观察到的有信息提示的复用与对不确定提示的持续探索。LEEPS将候选提示划分为利用和探索两个组合,并根据它们近期的非零奖励方差比例自适应分配rollout预算;它还利用表示空间邻居和历史rollout结果来优先考虑可能产生非零奖励方差的不确定提示,从而在不增加额外rollout的情况下让探索更具针对性。在6个数学推理基准上,LEEPS在两种模型规模下均取得最高平均得分,相比最强基线,Qwen2.5-Math-1.5B和7B的相对提升分别为2.6%和3.7%,且训练过程中通常收敛更快;在评估的3个OOD通用推理基准上,它在两种模型规模下也取得最高平均得分,且每训练步骤仅增加约2秒的在线采样开销。代码可在该https URL获取。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore--Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6\% and 3.7\% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at https://github.com/ShuangLiangX/LEEPS.
Comments15pages