不可学习,还是不可测量?关于RLVR中难度标签可靠性的研究
Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR
查看机构详情
- University of Dhaka(达卡大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究重新分析RLVR中的不可学习现象,发现困难提示以约三分之一的速率改善,难度标签不可靠,需更仔细的测量。
中文摘要 AI 辅助
基于可验证奖励的强化学习(RLVR)已成为在后训练阶段提升推理能力的重要方法。近期研究提出,一些困难提示(prompts)即使偶尔能产生正确解,仍难以通过训练得到改善。我们重新审视了这一“不可学习”现象,发现受影响的提示确实有所改进,其改进速率约为可学习提示的三分之一,而用于研究该现象的难度定义集合的可复现性远低于预期。这些难度标签是从有限数量的采样响应中估计得出的。跨随机种子(seeds)合并这些标签,不仅未能减少测量噪声,反而可能改变被选中的提示。我们开发了一个基于采样的框架来量化这种不稳定性,并确定需要多少评估量才能使难度分配可靠复现。我们还重新审视了先前提出的用于解释不可学习性的梯度相似性证据,并表明观察到的部分分离现象是因为困难提示提供的正确轨迹(rollouts)较少,导致其梯度估计不充分。匹配样本数量会削弱梯度差异,但不会消除它。总体而言,缓慢学习现象在我们的重新分析中依然存在,但用于定义该现象的提示和用于解释该现象的证据都需要更仔细的测量。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find that the affected prompts do improve, at roughly one third of the learnable rate, while the difficulty-defined set used to study them is much less reproducible than expected. These difficulty labels are estimated from a limited number of sampled responses. Combining them across seeds can further change which prompts are selected instead of simply reducing measurement noise. We develop a sampling-based framework for quantifying this instability and determining how much evaluation is required for difficulty assignments to reproduce reliably. We also revisit the gradient-similarity evidence proposed to explain unlearnability and show that part of the observed separation arises because difficult prompts provide fewer correct rollouts from which their gradients can be estimated. Matching this sample count weakens the gradient difference but does not remove it. Overall, the slow-learning phenomenon survives our reanalysis, while both the prompts used to define it and the evidence used to explain it require more careful measurement.