SHIFT-LLM:深度剪枝大语言模型中的分布偏移校正
SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs
浏览论文内容
中文总结 AI 辅助
SHIFT-LLM是一种无需训练的剪枝后校正框架,通过插入线性残差适配器缓解深度剪枝大语言模型的分布偏移,可在低样本无梯度下恢复准确率,在Llama-3.1-8B-Instruct上增益达15.7个点。
中文摘要 AI 辅助
深度剪枝通过移除整个Transformer模块来降低大语言模型的推理成本,但会破坏下游层所需的隐藏状态分布,导致显著的准确率损失。我们提出SHIFT-LLM,这是一种无需训练的剪枝后校正框架,在每个剪枝位置插入线性残差适配器(LRA)。每个LRA保留原始残差模块的恒等路径,并添加轻量级仿射残差校正。该校正通过在小型保留集上进行闭式最小二乘回归校准,无需梯度计算,以近似被剪枝模块产生的缺失残差更新。结合保留的恒等路径,最终的LRA输出近似原始模块产生的隐藏状态,从而缓解因层移除引入的分布不匹配,同时避免被移除模块的昂贵注意力和前馈计算。生成的LRA支持低秩分解和连续剪枝层之间的精确合并以实现额外压缩,且可与参数高效微调自然结合,进一步恢复仅微调剪枝模型之外的性能。在五个模型家族、六个层选择准则和七个零样本基准上的实验表明,SHIFT-LLM在大多数配置下能一致恢复深度剪枝导致的准确率损失,在Llama-3.1-8B-Instruct上实现了高达15.7个点的增益,且仅需数百个校准样本,无需梯度计算。
英文摘要
Depth pruning removes entire Transformer blocks to reduce the inference cost of large language models, but disrupts the hidden-state distributions expected by downstream layers, leading to significant accuracy loss. We introduce SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter (LRA) at each pruning site. Each LRA preserves the identity pathway of the original residual block and adds a lightweight affine residual correction. This correction is calibrated via closed-form least-squares regression on a small held-out set, without gradient computation, to approximate the missing residual update produced by the pruned block. Together with the preserved identity pathway, the resulting LRA output approximates the hidden state produced by the original block, thereby mitigating the distributional mismatch introduced by layer removal while avoiding the expensive attention and feed-forward computations of the removed blocks. The resulting LRAs support low-rank factorization and exact merging across consecutive pruned layers for additional compression, and combine naturally with parameter-efficient fine-tuning for further recovery beyond fine-tuning the pruned model alone. Experiments on five model families, six layer-selection criteria, and seven zero-shot benchmarks show that SHIFT-LLM consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct while requiring only a few hundred calibration samples and no gradient computation.
发表机构
- Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
机构由 AI 辅助整理,请以论文原文为准。