发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对在线策略蒸馏(OPD)的监督可靠性存疑的问题,提出无监督的在线策略自适应(OPSA)方法,在多个基准测试中较 OPD 和 Qwen3-1.7B 实现显著性能提升。
AI 中文摘要
在线策略蒸馏(OPD)提供了密集的 token 级监督,作为带可验证奖励的强化学习(RLVR)所产生的稀疏结果级优势的替代方案。然而,教师模型对学生生成的轨迹进行评分,这些轨迹本质上对教师而言是离线策略,因此其监督的可靠性以及学生改进的来源仍不清楚。我们对 OPD 训练期间的教师监督进行定量分析,发现存在大量噪声,且噪声的普遍性随教师模型规模的增大而增加。令人惊讶的是,学生策略对这种噪声并不敏感,无论是否保留或移除带噪声的监督,其收敛性能都相当。OPD 究竟是否能实现蒸馏?通过分析驱动其性能提升的因素,我们发现学习集中在低对数概率 token 上,且使用单个固定的负优势与教师提供的负优势性能相当。这表明 OPD 主要通过抑制低对数概率 token 发挥作用,而这一过程并不需要教师模型。这些发现促使我们提出了无监督方法——在线策略自适应(OPSA),该方法使用熵自适应负优势。它为高熵位置分配更强的学习信号,抑制尾部 token,并在头部 token 之间均匀重新分配概率质量。与基础模型 \texttt{Qwen3-1.7B} 相比,OPSA 在 AIME24 上的 Avg@32 提升了 35.41 个百分点,对应 263% 的相对增益,且在所有三个基准测试中 Pass@32 均提高了一倍以上;在 AIME24 上,其 Avg@32 较 OPD 高出 16.77 个百分点。跨模型家族和任务的大量实验与分析进一步证明了其有效性和可推广性。
英文摘要
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Comments23 pages, 14 figures