arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

压缩关键内容:神经元重要性与数据感知低秩近似用于语言模型压缩

Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression

Athanasios Ntovas, Alexandros Doumanoglou, Petros Drakoulis, Dimitris Zarpalas

arXiv 2607.18284首次发表:更新:

发表机构

Information Technologies Institute (ITI), Centre for Research and Technology HELLAS (CERTH)(信息技术研究所(ITI),希腊研究与技术中心(CERTH))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言模型压缩问题,结合神经元重要性和数据感知低秩近似方法,提出动态压缩率分配算法,实验表明该方法在高压缩率下性能与或优于现有技术。

AI 中文摘要

为在其领域表现出色,大型语言模型由数十亿参数组成,这导致巨大内存需求,限制了其在资源受限环境中的应用。为解决神经网络压缩问题,奇异值分解在矩阵压缩中起关键作用。此前工作分别从参数重要性或每层功能等效性角度关注神经网络权重矩阵的低秩近似。本文研究将这两个视角结合在一个目标中的方法的有效性。同时,影响压缩质量的一个重要方面是压缩率在各层和神经网络参数间的分布。早期工作大多均匀分配压缩率或依赖计算昂贵的启发式搜索。本文提出一种增强且计算高效的动态压缩率分配算法。实验结果支持了该方法的有效性,其性能与之前的最先进方法相当或显著更好,尤其是在高压缩率下。

英文摘要

To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression Singular Value Decomposition (SVD) has played a key role as a fundamental component for matrix compression through decomposition. To minimize compression error and to maximize the efficacy of the compressed model on the downstream tasks previous works focused on low-rank approximation of the NN's weight matrices either from the perspective of parameter importance or per-layer functional equivalence. While previous works studied the aforementioned perspectives in isolation in this work we are investigating the effectiveness of an approach that combines ideas from these two perspectives in a single objective. In parallel to this an important aspect that affects the compression quality is the distribution of the compression rate across layers and NN parameters. Earlier works mostly considered distributing the compression rate uniformly across layers and network weights or relied on computationally expensive heuristic search. Contrary to them in this work we propose an enhanced and computationally efficient algorithm for dynamic compression rate allocation. Experimental results support the efficacy of the proposed approach which performs on par or substantially better than the previous state-of-the-art especially under high compression ratios.

Journal refEEE Access, vol. 14, pp. 6106-6120, 2026

DOI:10.1109/ACCESS.2026.3653132

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑