将技能蒸馏为权重而非提示:将抽象技能作为在线自蒸馏的特权信号
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
浏览论文内容
中文总结 AI 辅助
该研究针对强化学习中组相对奖励失效的问题,提出SKALD框架,通过将抽象技能蒸馏为模型权重,在五个数学基准上提升了不同规模模型的avg@8指标。
中文摘要 AI 辅助
带有可验证奖励的强化学习在推理组全部正确或全部错误时,无法产生组相对信号,这类情况在实验中占比达63.0%-68.0%。我们提出SKALD(Skill-Anchored Latent Distillation,技能锚定潜在蒸馏),这是一个在线自蒸馏框架,使用同一Qwen3-Base模型的两种上下文视图:仅含问题的学生模型,以及以抽象、显式答案过滤的技能卡为条件的教师模型。学生模型在自身前缀上训练,将技能诱导的优势转移到共享参数中,测试时无需特权输入。为稳定上下文诱导的分布不匹配,SKALD采用退火指数倾斜目标,对学生模型概率极低的教师偏好令牌进行降权;随着倾斜消失,该目标收敛到教师交叉熵并恢复前向KL学生梯度。一个经验门控机制仅在可验证推理估计出正教师优势时激活蒸馏。在五个保留的数学基准上,SKALD在0.6B、1.7B、4B参数规模下,整体avg@8较GRPO分别提升+2.46、+4.85、+12.01;在1.7B参数规模下,仅零方差蒸馏可恢复全部增益的84.7%,而SKALD较计算量匹配的GRPO仍提升+4.06,且超出上下文技能暴露的效果+3.77。这些结果表明,在组相对奖励变得无信息时,抽象技能能提供密集监督信号。
英文摘要
Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Skill-Anchored Latent Distillation), an on-policy self-distillation framework that uses two context views of the same Qwen3-Base model: a question-only student and a teacher conditioned on an abstract, explicit-answer-filtered skill card. The student is trained on its own prefixes, transferring the skill-induced advantage into shared parameters without privileged input at test time. To stabilize context-induced distribution mismatch, SKALD employs an annealed exponentially tilted objective that downweights teacher-preferred tokens with very low student likelihood; as the tilt vanishes, it converges to teacher cross-entropy and recovers the forward-KL student gradient. An empirical gate activates distillation only when verified rollouts estimate a positive teacher advantage. Across five held-out mathematics benchmarks, SKALD improves overall avg@8 over GRPO by +2.46, +4.85, and +12.01 at 0.6B, 1.7B, and 4B, respectively. At 1.7B, zero-variance-only distillation recovers 84.7% of the full gain, while SKALD remains +4.06 above FLOP-matched GRPO and exceeds contextual skill exposure by +3.77. These results show that abstract skills provide dense supervision where group-relative rewards become uninformative.