arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReCal:针对策略蒸馏恢复的结构化剪枝校准

ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery

Houcheng Jiang, Mao Zheng, Mingyang Song, Qiyong Zhong, Jie Sun, Tianyu Zhang, Junfeng Fang

arXiv 2610.11332首次发表:更新:

发表机构

Zhongguancun Academy; Tencent; University of Science and Technology of China; National University of Singapore(中关村学院; 腾讯; 中国科学技术大学; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对结构化剪枝导致推理语言模型的策略蒸馏(OPD)恢复受阻的问题,提出即插即用方法RECAL,通过调整剪枝前的校准提升OPD恢复效果,在AIME等数据集上取得显著性能增益。

AI 中文摘要

结构化剪枝可降低推理语言模型的部署成本,但由此产生的性能下降会阻碍后续的策略蒸馏(OPD)恢复。由于OPD依赖学生模型生成的轨迹,离线蒸馏后仍存在的剪枝损伤会限制其有效性。我们提出RECAL(Recovery-Aware Calibration,即感知恢复的校准),这是一种简单的即插即用方法,通过在剪枝前调整校准来提升OPD恢复效果。RECAL利用未剪枝教师模型与剪枝探针之间的前向KL散度,识别被剪枝破坏的教师支持预测,随后重新加权校准统计量,以引导现有剪枝准则保留这些预测。在多个模型和剪枝方法上,RECAL经OPD后持续提升数学推理能力,在AIME数据集上最高提升16.7个百分点,多数代码生成对比中也有改善。进一步分析显示,RECAL减少了受影响严重的token处的残留损伤,并建立了可在恢复过程中持续的性能优势。这些结果表明,感知恢复的校准对提升剪枝后推理模型的策略蒸馏恢复具有重要价值。

英文摘要

Structured pruning reduces the deployment cost of reasoning language models, but the resulting capability degradation can hinder subsequent on-policy distillation (OPD) recovery. Because OPD relies on student-generated trajectories, pruning damage that persists after offline distillation can limit its effectiveness. We propose RECAL, Recovery-Aware Calibration, a simple plug-and-play approach that improves OPD recovery by adjusting calibration before pruning. RECAL uses forward KL between an unpruned teacher and a pruned probe to identify teacher-supported predictions disrupted by pruning, then reweights calibration statistics to guide existing pruning criteria toward preserving these predictions. Across multiple models and pruning methods, RECAL consistently improves mathematical reasoning after OPD, achieving gains of up to 16.7 percentage points on AIME, alongside improvements in most code-generation comparisons. Further analysis shows that RECAL reduces residual damage at heavily affected tokens and establishes performance advantages that persist through recovery. These results demonstrate the value of recovery-aware calibration for improving on-policy distillation recovery of pruned reasoning models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑