可微分位宽:通过SVD协同优化剪枝与量化实现超高效LLM压缩
Differentiable Bit-Widths: Co-optimizing Pruning and Quantization via SVD for Ultra-Efficient LLM Compression
- Ajou University(亚洲大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出可微分位宽方法,在统一框架中协同优化剪枝与量化,通过SVD学习组件级位宽,实现超高效LLM压缩,优于两阶段基线。
AI中文摘要:
基于SVD的剪枝和量化最近成为超高效压缩大型语言模型的一种有前景的策略。在这些方法中,压缩分两个阶段进行:首先截断组件,然后对剩余组件进行量化。尽管这种解耦流程受益于剪枝和量化,但它需要对每个阶段进行单独优化,无法充分利用两者之间的平衡,这可能导致在激进压缩下性能次优。为解决这一限制,我们提出了一种新的LLM压缩方法,在统一框架中协同优化剪枝和量化。我们的关键思想是一种可微分的组件级位宽学习方法,允许较不重要的组件被分配0位精度并被剪枝掉。值得注意的是,即使在为超高效设计的极端量化设置(1.61位)下,我们的方法也优于两阶段基线。代码:此https URL。
英文摘要:
SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings ($1.61$ bits) designed for ultra-efficiency. Code: https://github.com/MMAI-Laboratory/DBW.