AI 中文总结
本研究探讨在线策略蒸馏在弱到强等设置中的缩放性质,发现早期训练存在有用转移阶段,学生峰值分数可超教师,并拟合幂律揭示教师规模与转移效果的关系。
AI 中文摘要
强化学习(RL)能够在大语言模型(LLMs)中引发显著的推理能力,但这种能力在多大程度上跨模型规模转移,以及转移速度有多快,仍不清楚。我们研究了在线策略蒸馏(OPD)在弱到强、同基座和强到弱教师-学生设置中的缩放性质。我们发现,早期OPD训练动态统一表现出一种规律的“有用转移”阶段,在此阶段,留出准确率(黄金分数,$G$)近似线性地随$d=\sqrt{\mathrm{KL}(\pi_\theta \Vert \pi_{\mathrm{ref}})}$(即学生初始化时token级反向KL散度的平方根)增加。在每个观察到的弱到强配对中,学生的峰值黄金分数超过其教师自身的分数,因此紧凑的RL专家可以通过OPD将能力转移给更大的学生。为了估计OPD结果,我们拟合了幂律,用于描述$G_{\mathrm{peak}}$以及有用转移阶段的斜率如何随学生和教师参数数量以及教师黄金分数缩放。这些定律表明,峰值黄金分数仅随教师规模增加到大约学生规模为止,并且在匹配的黄金分数下,较小的教师转移效果更好,因此教师的分数本身并不能定义其监督价值。我们还研究了两种OPD变体的缩放效应:自举弱到强OPD,以及在线策略监督的程度。
英文摘要
*Reinforcement learning (RL)* can induce substantial reasoning capabilities in large language models (LLMs), but how much of this capability transfers across model scales, and how quickly, remains unclear. We study the scaling properties of *on-policy distillation (OPD)* across *weak-to-strong*, *same-base*, and *strong-to-weak* teacher--student setups. We find that early OPD training dynamics uniformly exhibit a regular *useful-transfer* regime, in which held-out accuracy (the *gold score*, $G$) rises approximately linearly in $d=\sqrt{\mathrm{KL}(π_θ\Vert π_{\mathrm{ref}})}$, the square root of token-level reverse KL divergence from the student initialization. In every observed weak-to-strong pair, the student's peak gold score exceeds its teacher's own, so a compact RL expert can transfer capability to a much larger student via OPD. To estimate OPD outcomes, we fit *power laws* for how $G_{\mathrm{peak}}$ and the slope of the useful-transfer regime scale with student and teacher parameter counts and with teacher gold score. These laws show that peak gold score improves with teacher scale only up to roughly the student's scale, and that at a matched gold score smaller teachers transfer better, so a teacher's score alone does not define its supervision value. We also study the scaling effects of two OPD variants, bootstrapping weak-to-strong OPD, and the degree of on-policy supervision.
Comments35 pages, 20 figures