arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17941cs.LGcs.AIcs.CL

基于图结构在线难度估计的高效RLVR调度

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对RLVR中探索预算分配低效问题,提出即插即用的图结构在线难度估计器,可跨样本共享反馈、缓解冷启动与过时问题,集成后实现难度自适应探索,在多模型与基准上性能更优。

中文摘要 AI 辅助

带可验证奖励的强化学习(RLVR)可提升大语言模型的推理能力,但依赖代价高昂的rollout探索。为不同难度的样本分配相同探索预算效率低下:简单样本可能获得冗余rollout,而难以学习的样本可能获得过少探索。现有自适应调度器通过基于课程的样本选择或基于估计样本难度的非均匀rollout分配解决这种不匹配,但获取可靠的在线难度估计仍具挑战性:专用探测会增加大量生成开销,而基于历史的估计器面临无初始观测的冷启动和反馈过时问题,且通常忽略样本间的关系。为解决这些局限,我们提出一种即插即用的基于图的在线难度估计器,其跨相关样本共享rollout反馈并持续更新它们的难度估计,无需专用探测即可缓解冷启动和过时问题。具体而言,我们首先基于语义和推理相似性构建感知难度的样本图;基于该图,引入潜在难度状态并使用Potts先验鼓励相邻样本共享相同状态;随后采用状态级Beta-Binomial模型聚合与每个状态相关的rollout结果;最后,使用在线平均场变分算法在新反馈到达时持续更新潜在状态分配和状态级难度。我们的框架可集成到样本选择和rollout分配调度器中,实现无需专用探测的难度自适应探索。在多个基础模型、RL调度器和基准上的实验表明,我们的框架取得了更好的性能。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.

↑