arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大语言模型的输出感知残差流剪枝

Output-aware Residual Stream Pruning for Large Language Models

Chayne Thrash, Kevin Chen, Soheil Kolouri

arXiv 2609.35579首次发表:更新:

发表机构

Vanderbilt University(范德比尔特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种敏感性感知的残差流剪枝方法,通过输出KL散度的二阶近似和谱上界,将子空间选择简化为特征分解,在多个指令微调模型上优于仅基于激活的剪枝,提升困惑度和下游任务性能。

AI 中文摘要

残差流剪枝方法通过缩小模型的隐藏维度来降低推理成本,但现有方法通常通过最小化激活重建误差来选择这些维度。这一标准隐含地认为所有扰动方向同等重要,忽略了下游层的敏感性。我们提出了一种对敏感性感知的残差流剪枝方法,直接考虑这种方向相关的敏感性。利用输出KL散度的二阶近似,我们通过激活协方差和模型输出的局部敏感性来刻画残差流扰动的影响。由此产生的子空间选择目标耦合了这两个量,但难以直接优化。我们推导出一个可处理的谱上界,将子空间选择简化为对敏感性加权协方差矩阵的特征分解,保留了基于旋转的剪枝方法的效率和结构简单性。在多个指令微调语言模型家族中,我们的方法相对于仅基于激活的剪枝,持续降低了校准KL散度,并在多种压缩水平下改善了困惑度和下游任务性能。我们的结果表明,仅保留激活能量不足以进行残差流剪枝,而显式考虑扰动如何传播到模型输出,为选择要移除的维度提供了更有效的标准。

英文摘要

Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities, but is difficult to optimize directly. We derive a tractable spectral upper bound that reduces subspace selection to an eigendecomposition of a sensitivity-weighted covariance matrix, retaining the efficiency and structural simplicity of rotation-based pruning methods. Across several instruction-tuned language model families, our method consistently reduces calibration KL divergence relative to activation-only pruning and improves perplexity and downstream task performance over a range of compression levels. Our results show that preserving activation energy alone is insufficient for residual-stream pruning, and that explicitly accounting for how perturbations propagate to the model output provides a more effective criterion for selecting dimensions to remove.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑