发表机构
School of Computer Science and Technology, Guangdong University of Technology; College of Information and Artificial Intelligence, Yangzhou University(广东工业大学计算机科学与技术学院; 扬州大学信息与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对Transformer模型部署的高资源消耗问题,提出基于LoRA低秩梯度的REP-LIE剪枝方法,结合稳定性分数迭代剪枝,在LLaMA-7B等模型上实现高效剪枝并保持竞争力性能。
AI 中文摘要
随着基于Transformer架构的大规模预训练语言模型快速发展,其高昂的计算与内存成本已成为部署的主要障碍,尤其在资源受限环境中。传统剪枝方法通常依赖基于全梯度的重要性估计,且需预先对模型进行微调以获得满意性能,该过程常导致难以承受的资源消耗。本文提出REP-LIE,一种在微调过程中实现资源高效剪枝的新方法。REP-LIE利用LoRA低秩矩阵的梯度估计权重重要性,无需全梯度计算。为解决重要性估计中的固有随机性,引入稳定性分数,作为迭代剪枝不重要模型参数的依据。剪枝后的模型通过轻量级更新进一步微调,无需在微调过程中进行全参数优化。在中等规模编码器模型及大规模生成模型(LLaMA-7B与Mistral-7B)上开展的大量实验表明,REP-LIE与现有方法相比仍能达到有竞争力的性能。
英文摘要
With the rapid development of large-scale pre-trained language models based on Transformer architectures, their high computational and memory costs have become a major obstacle to deployment, especially in resource-constrained environments. Traditional pruning methods typically depend on full gradient-based importance estimation, and they necessitate prior finetuning of the model to achieve satisfactory performance. This process often results in intolerable resource consumption. This paper proposes REP-LIE, a new approach to enable resource-efficient pruning during the process of finetuning. REP-LIE leverages the gradients of LoRA low-rank matrices to estimate the importance of weights without requiring full gradient computation. To address the inherent randomness in importance estimation, a stability score is introduced, serving as the basis for iterative pruning of unimportant model parameters. The pruned model is further finetuned through lightweight updates, eliminating the need for full-parameter optimization in the process of finetuning. Extensive experiments on both medium-scale encoder models and large-scale generative models (LLaMA-7B and Mistral-7B) demonstrate that REP-LIE still achieves competitive performance compared to existing approaches.
Comments15 pages, 9 figures
Journal refIEEE Transactions on Emerging Topics in Computational Intelligence, 2026
DOI:10.1109/TETCI.2026.3727362