arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36120cs.LGcs.AI

ThinQuant:用于大语言模型权重与激活量化的可扩展旋转学习

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

Mehdi Makni, Ryan Lucas, Rahul Mazumder

首次发表
浏览论文内容

中文总结 AI 辅助

ThinQuant通过数据选择和精确优化约简,实现了高效的无梯度旋转学习,在低比特量化中匹配甚至超越现有方法,并首次将旋转校准扩展到405B模型。

中文摘要 AI 辅助

学习旋转在大语言模型的低比特权重和激活量化中发挥着重要作用,它通过平滑激活分布中的异常值来实现量化。最先进的方法包括基于梯度的过程(如SpinQuant)和计算上更友好的无梯度方法(如DartQuant),但两者都难以扩展到最大的架构。为了解决无梯度旋转学习中的计算瓶颈,我们引入了两个提高效率的思路:(i)一种数据选择程序,减少了所需的校准数据点数量;(ii)在此缩减的校准集上对相关优化进行精确约简。我们的数据选择程序利用了激活凸包的几何结构。利用这一思路,我们表明,与最先进的基于旋转的方法相比,精心选择的校准集其激活数量少几个数量级,却能在低比特量化设置中匹配其性能。在这种极端数据效率下,选中的激活张成一个$r$维子空间($r<d$),使得对$d\ imes d$旋转的优化等价于在Stiefel流形上优化一个$d\ imes r$矩阵。我们使用一种高效的ADMM算法来解决这个约简问题,该算法在每一步迭代地采用薄矩阵更新,因此得名ThinQuant。对于Llama-3-70B的W4A4KV4量化,ThinQuant在不到12分钟内完成整个旋转校准,并达到WikiText-2困惑度5.63,而DartQuant需要111分钟,困惑度为7.55。与SpinQuant和DartQuant不同,ThinQuant还能在单个H200 GPU上扩展到Llama-3.1-405B,仅用2小时多一点完成旋转校准,在W4A4下达到WikiText-2困惑度2.97,而GPTAQ+QuaRoT为3.48。

英文摘要

Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structure of the convex hull of the activations. Using this idea, we show that a carefully selected calibration set with several orders of magnitude fewer activations than state-of-the-art rotation-based methods can match their performance in low-bit quantization settings. Under this extreme data efficiency, the selected activations span an $r$-dimensional subspace with $r<d$, making optimization over a $d\times d$ rotation equivalent to optimizing a $d\times r$ matrix on the Stiefel manifold. We solve this reduced problem using an efficient ADMM algorithm that iteratively employs thin matrix updates at every step, hence the name ThinQuant. For Llama-3-70B with W4A4KV4 quantization, ThinQuant completes the entire rotation calibration in under 12 minutes and achieves a WikiText-2 perplexity of 5.63, compared with 7.55 for DartQuant, which requires 111 minutes. Unlike SpinQuant and DartQuant, ThinQuant also scales to Llama-3.1-405B on a single H200 GPU, completing rotation calibration in just over 2 hours and achieving WikiText-2 perplexity of 2.97 at W4A4, compared with 3.48 for GPTAQ+QuaRoT.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑