arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09639cs.CL

在策略蒸馏传授新技能而非新知识

On-Policy Distillation Teaches New Skills but Not New Knowledge

发表机构香港科技大学
查看机构详情
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yixuan Tang, Yi Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过受控合成框架和多种模型实验,发现反向KL在策略蒸馏主要传递组合推理技能而非新事实知识,而前向KL可恢复事实传递,表明该方法教会模型组织已有知识而非扩展参数知识。

中文摘要 AI 辅助

在策略蒸馏(OPD)增强了语言模型的推理能力,然而学生模型是否获得了新的事实知识或多步推理的组合技能仍属未知。我们通过一个受控的合成框架将这些能力分离,该框架衡量学生模型的初始能力,并独立控制教师模型额外的事实、组合技能或两者兼有。在来自三个家族的四种模型上,反向KL的OPD可靠地跨未见推理结构传递组合技能,但传递的事实知识极少。解耦蒸馏配方揭示了这种不对称性的来源:用前向KL替换反向KL恢复了事实传递,而学生模型的滚动生成则特别改善了多步推理的执行。在近期事实问答和竞赛数学上的实验显示,在反向KL的OPD下存在类似的不对称性,产生了显著的推理增益而无事实记忆扩展。综合来看,这些结果表明在策略蒸馏并未扩展模型的参数化知识,而是教会它组织和组合其已拥有的知识。

英文摘要

On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.

↑