发表机构
ByteDance; Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences(字节跳动; 中国科学院软件研究所; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究结构化剪枝在语言模型自由形式生成任务中存在的问题,提出ShortOPD方法,通过短到长的策略蒸馏,检测重复后缀,合理分配展开预算,有效提升压缩模型分数,减少训练时间和展开令牌,推动结构化剪枝接近可部署的生成质量。
AI 中文摘要
结构化剪枝是一种对硬件友好的语言模型压缩方式,但大多在多项选择识别任务中得到验证,相同的压缩检查点在实际部署所需的自由形式生成任务中可能会崩溃。本文通过两项观察发现了这种差距。首先,贪心的\textsc{pass}@$1$在压缩后几乎消失,但\textsc{pass}@$k$在重复采样下能大幅恢复。其次,可恢复机制主要因后缀重复而失败。因此,恢复应在压缩模型自身的策略状态上进行密集的令牌级监督训练,策略蒸馏(OPD)通过将预压缩模型用作冻结教师来提供这种监督。然而,长时间的策略展开会将早期恢复预算花费在低信息重复后缀上,延迟损失下降。为缓解这种浪费,本文提出了\textbf{\shortopd},一种短到长的OPD调度,它能检测教师确认的重复后缀,将幸存的前缀视为每次展开的有效长度,并将未来的展开预算分配给策略当前可使用的有效长度。在数学、代码和开放式生成任务中,\shortopd\将压缩模型的分数提高到未恢复值的约$9$倍,以及标准恢复方法(无知识蒸馏的监督微调、知识蒸馏和序列知识蒸馏)的$1.6$ - $4.4$倍,并且在两点内匹配固定的$8192$令牌展开范围,使用四分之一的训练时间($8.5$小时对$35.9$小时)和减少$71\%$的展开令牌。希望该方法有助于使结构化剪枝超越在困惑度和多项选择基准上的微小收益,更接近可部署的生成质量。
英文摘要
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires. Two observations trace this gap. First, greedy \textsc{pass}@$1$ nearly vanishes after compression, yet \textsc{pass}@$k$ recovers substantially under repeated sampling: useful generations are demoted, not erased. Second, the recoverable regime fails mainly through suffix repetition. Recovery should therefore train on the compressed model's own on-policy states with dense token-level supervision, which On-Policy Distillation (OPD) provides by reusing the pre-compression model as a frozen teacher. However, long on-policy rollouts spend early recovery budget on low-information repetitive suffixes, delaying loss descent. To mitigate this waste, we propose \textbf{\shortopd}, a short-to-long OPD schedule that detects teacher-confirmed repetitive suffixes, treats the surviving prefix as each rollout's effective length, and allocates future rollout budgets to the effective lengths the policy can currently use. Across math, code, and open-ended generation, \shortopd\ raises the compressed model's score to about $9\times$ its unrecovered value and $1.6$--$4.4\times$ standard recovery recipes (SFT w/o KD, KD, and SeqKD), and it matches a fixed $8192$-token rollout horizon within two points using a quarter of the training time ($8.5$ vs.\ $35.9$ hours) and $71\%$ fewer rollout tokens. We hope this recipe helps move structured pruning beyond marginal gains on perplexity and multiple-choice benchmarks, a step closer to deployment-ready generation quality.