发表机构
New York University; Amazon(纽约大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对模型-技能协同进化中技能蒸馏的选择问题,提出SGUID方法,选定紧凑技能子集蒸馏可提升性能,还能实现稳定协同进化,在Qwen系列模型上取得了显著效果。
AI 中文摘要
技能是在推理阶段添加的可复用过程性指导,可大幅提升大语言模型(LLM)的下游性能(Li等人,2026)。现有研究通过语义相关性从技能库中检索技能,将其用作推理时的补丁或用于模型蒸馏,但各技能的个体效用在很大程度上被忽视。我们首先表明,在以技能条件策略作为教师的在线策略蒸馏中,不到25%的检索技能能提供有用的蒸馏信号。随后我们提出SGUID,一种用于选择紧凑技能子集进行蒸馏的方法:仅保留在训练过程中持续产生有效学习信号的技能,再对所选技能进行蒸馏以得到更优模型。我们的结果显示并非所有技能都值得蒸馏:在Olmo和Qwen系列的四个模型中,对6个选定技能进行蒸馏,在其中三个模型上的平均avg@12指标与全库蒸馏相当或更优;在第二轮蒸馏3个新选定技能后,四个模型均达到该效果,而全库规模是选定技能库的11倍之多。重要的是,SGUID支持稳定的模型-技能协同进化:每轮蒸馏后,会从更新后模型的 rollout 结果中整理出新的候选技能库,SGUID会选择接下来要内化的技能;在第二轮中,该循环选定3个新技能,将Qwen3-8B的性能从64.3%提升至66.3%。选择步骤对稳定性至关重要:在Qwen3-4B上,直接用未过滤的技能更新模型会导致性能下降,包括HMMT25上0.3个百分点的降幅,而SGUID在第一轮后将HMMT25提升0.5个点,第二轮后提升1.1个点。这些结果表明,技能选择是实现稳定模型-技能协同进化的关键机制。
英文摘要
Skills, reusable procedural guidance added at inference, can substantially improve LLM downstream performance (Li et al., 2026). Prior work retrieves skills from a bank by semantic relevance, then uses them as inference-time patches or for model distillation. The individual utility of each skill, however, is largely neglected. We first show that, in on-policy distillation where skill-conditioned policies serve as teachers, fewer than 25% of retrieved skills provide useful distillation signals. We then propose SGUID, a method for selecting a compact subset of skills for distillation. SGUID retains a skill only if it consistently yields effective learning signals during training. The selected skills are then distilled to produce a better model. Our results show that not all skills are worth distilling. Across four models from the Olmo and Qwen families, distilling 6 selected skills matches or exceeds full-bank distillation in mean avg@12 on three of the four models, and on all four after a second round that distills 3 newly selected skills, while the full banks are up to 11x larger. Importantly, SGUID supports stable model-skill co-evolution: after a distillation round, a new candidate bank is curated from the updated model's rollouts, and SGUID selects which skills to internalize next. In the second round, this loop selects 3 new skills and improves Qwen3-8B from 64.3% to 66.3%. The selection step is essential for stability: on Qwen3-4B, naively updating the model with unfiltered skills degrades performance, including a 0.3 percentage point drop on HMMT25, whereas SGUID improves HMMT25 by 0.5 points after the first round and 1.1 points after the second. These results identify skill selection as the key mechanism for stable model-skill co-evolution.