arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

免前向传播的基于权重冗余的大语言模型深度剪枝

Forward-Free LLM Depth Pruning via Weight Redundancy

Vincent-Daniel Yun, Woosang Lim

arXiv 2609.09883首次发表:更新:

发表机构

University of Southern California; Seoul National University(南加州大学; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出免前向传播的权重冗余剪枝方法,利用检查点权重估计层间冗余选择Transformer块,无需校准数据,性能优于现有免前向传播方法并接近基于激活的方法。

AI 中文摘要

深度剪枝通过移除完整的Transformer块来降低大语言模型(LLM)的推理成本。基于激活的方法通过在前向传播过程中对校准数据收集隐藏状态,而现有的免前向传播方法则单独评估每个Transformer块,不测量块之间的相似性。我们提出了权重冗余剪枝(WRP),一种免前向传播的深度剪枝方法,它从检查点权重中估计层间冗余以选择块,无需校准数据或模型前向传播。WRP比较各层之间的注意力输出和MLP下投影权重,并结合它们的成对相似性与相对投影尺度信息。由此产生的全对相似性矩阵指导层分组和块选择。在多种剪枝设置、模型家族和下游任务中,WRP始终优于现有的免前向传播幅度剪枝方法,并接近基于激活的方法的性能。

英文摘要

Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, while existing forward-free methods score each Transformer block separately without measuring similarity between blocks. We propose Weight-Redundancy Pruning (WRP), a forward-free depth-pruning method that estimates inter-layer redundancy from checkpoint weights to select blocks without calibration data or model forward passes. WRP compares attention output and MLP down-projection weights across layers and combines their pairwise similarities with relative projection-scale information. The resulting all-pairs similarity matrix guides layer grouping and block selection. Across multiple pruning settings, model families, and downstream tasks, WRP consistently outperforms existing forward-free magnitude pruning and approaches the performance of activation-based methods.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑