AI 中文总结
本研究提出无训练补偿框架COEC,将其应用于Llama-3等模型的结构化剪枝,可提升剪枝后模型的困惑度与零样本准确率,在高稀疏度下增益更显著。
AI 中文摘要
结构化剪枝通过移除权重列来减小大语言模型(LLM)的规模并降低推理成本,但由此产生的输出误差会降低模型准确率。现有的无训练补偿方法会在保留权重的输出侧使用加性偏置或单一正交旋转,这些修正保留了其输入奇异帧不变,因此限制了保留权重在列移除后的适应能力。我们提出COEC(Calibrated Orthogonal-Equivalence Compensation,校准正交等价补偿),这是一种无训练补偿框架,它对保留权重应用交替的左、右正交旋转。右旋转在简化的斯蒂费尔流形上进行优化,同时使用广义交叉验证重新缩放奇异值以选择各层的正则化强度。COEC进一步调整校准格拉姆矩阵以降低高能量激活方向的主导地位,并引入对齐惩罚项以保留相邻注意力组件之间的几何关系。这些组件使用小型校准集的二阶统计量,且无需对LLM进行反向传播或重新训练模型参数。COEC独立于列剪枝准则,可应用于多种结构化剪枝方法。在Llama-3、Llama-3.1和Qwen2.5模型系列上针对多个结构化稀疏度水平开展的实验表明,与现有补偿方法相比,COEC在所有模型上均提升了困惑度,且在大多数设置下提升了零样本准确率,在更高稀疏度下增益更大。这些结果表明,剪枝后补偿可恢复部分因列移除而损失的性能。
英文摘要
Structured pruning reduces the size and inference cost of large language models (LLMs) by removing weight columns, but the resulting output error can degrade accuracy. Existing training-free compensation methods use an additive bias or a single orthogonal rotation on the output side of the retained weight. These corrections leave its input singular frame unchanged and therefore limit how the retained weight can adapt after column removal. We propose COEC (Calibrated Orthogonal-Equivalence Compensation), a training-free compensation framework that applies alternating left and right orthogonal rotations to the retained weight. The right rotation is optimized on a reduced Stiefel manifold, while singular values are rescaled using generalized cross-validation to select the regularization strength for each layer. COEC further tempers the calibration Gram matrix to reduce the dominance of high-energy activation directions and introduces an alignment penalty that preserves the geometric relation between adjacent attention projections.All components use second-order statistics from a small calibration set and require neither backpropagation through the LLM nor retraining of the model parameters. COEC is independent of the column pruning criterion and can be applied to multiple structured pruning methods. Experiments on the Llama-3, Llama-3.1, and Qwen2.5 model families across multiple structured sparsity levels show that COEC improves perplexity on every model and zero-shot accuracy in most settings over existing compensation methods, with larger gains at higher sparsity. These results show that post-pruning compensation can recover part of the performance lost to column removal.