arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19213cs.LGcs.AI

面向高效LLM压缩的逐层课程学习

Layer-wise Curriculum Learning for Efficient LLM Compression

Donggeon Lee, Dooyeon Na, Seungmin Oh, Jongbin Ryu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出逐层课程学习压缩LLM,通过分段知识迁移与特征缓存,在BERT和GPT-2上减少超50%GPU内存和训练时间,并在LLaMA和Qwen上超越其他剪枝方法。

中文摘要 AI 辅助

在本文中,我们引入了逐层课程学习以实现高效的LLM压缩。所提出的方法促进了从教师模型到学生模型的知识迁移,利用了一种从较简单的优化任务开始并逐步处理较难任务的课程学习方法。为了在LLM压缩中采用逐层学习,我们将整个模型划分为多个由层组成的段,从而为LLMs实现更具计算效率的知识迁移。基于我们对累积误差现象的理论分析,逐层课程学习加速了收敛,同时稳定了知识迁移过程。此外,我们提出了一种具有多线程策略的特征缓存方法,以有效解决跨层的特征错位问题,最大化GPU利用率。因此,我们的方法展现出先进的模型压缩性能,以及在最小化内存使用和缩短训练时间方面的高计算效率。在多个数据集上的实验表明,所提出的方法在BERT和GPT-2上实现了最先进的性能,同时将GPU内存使用和训练时间减少了超过50%。此外,在相同的训练时间下,它在LLaMA系列和Qwen模型上优于其他剪枝方法,且GPU内存占用更低。

英文摘要

In this paper, we introduce layer-wise curriculum learning for efficient LLM compression. The proposed method facilitates the knowledge transfer from the teacher model to the student model, utilizing a curriculum learning approach that begins with easier optimization tasks and progressively tackles harder ones. In order to adopt the layer-wise learning in LLM compression, we partition the whole model into multiple segments consisting of layers, thereby enabling more computationally efficient knowledge transfer for LLMs. Based on our theoretical analysis of cumulative error phenomenon, layer-wise curriculum learning accelerates convergence while stabilizing the knowledge transfer process. In addition, we present a feature caching method with a multi-threading strategy to efficiently address feature misalignment across layers, maximizing GPU utilization. Consequently, our method exhibits advanced model compression performance, as well as high computational efficiency in terms of minimized memory usage and short training hours. Experiments on multiple datasets show that the proposed method achieves state-of-the-art performance while reducing GPU memory usage and training hours by more than 50\% on BERT and GPT-2. Moreover, it outperforms the other pruning methods on LLaMA-family and Qwen models under the same training hours, with a lower GPU memory footprint.

发表机构

  • Ajou University(亚洲大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑