arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BaKron:基于克罗内克因式分解海森矩阵的高效量化方法

FastKron: Efficient Quantization with Kronecker-Factored Hessians

Johann Birnick, Rayan Saab

arXiv 2608.06291首次发表:更新:

发表机构

University of California San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BaKron是一种高效神经网络量化算法,通过结合反对角并行与递归分治结构,将计算量从O(m²n²)降至O(mn(m+n)),兼具GPTQ的缩放效率与更丰富的曲率信息,且对基础量化器和海森估计器模块化。

AI 中文摘要

我们加速一类神经网络量化算法,这类算法的几何由海森矩阵的任意克罗内克因式分解近似所决定。GPTQ风格的自适应舍入通常使用从输入激活中得到的单侧信息,而双侧克罗内克因式分解海森矩阵近似还能捕捉输出坐标间的相关性,但直接在向量化权重域应用GPTQ计算成本很高。基于BoA和YAQA所用的双侧自适应舍入公式,我们提出BaKron,一种结合反对角并行性与递归分治结构的高效求解器。对于m×n的权重矩阵,BaKron使用O(m+n)的顺序步骤,同时将总计算量从O(m²n²)降至O(mn(m+n)),因此它达到了GPTQ的立方级缩放,同时利用了更丰富的曲率信息。此外,BaKron对于基础量化器和海森估计器均具有模块化特性。我们还提供了实际基准测试,考虑了BaKron可调用的一系列海森矩阵,找到一种计算这些海森矩阵的高效技术,并对该算法进行了实验评估。

英文摘要

We accelerate a family of algorithms for neural network quantization which utilize a two-sided version of the GPTQ/LDLQ algorithm. Standard GPTQ-style adaptive rounding uses one-sided correlation information derived from input activations. A natural two-sided extension can additionally capture correlations across output channels. It utilizes a general Kronecker-factored approximation of the weight matrix curvature. This approach has been used in BoA and YAQA. But making a concrete algorithmic implementation of this two-sided GPTQ variant is nontrivial. BoA uses a large number of sequential steps, while YAQA improves over the sequential depth but still has quartic total cost. We introduce FastKron, an efficient algorithmic implementation that combines anti-diagonal parallelism with a recursive divide-and-conquer construction. For an $m\times n$ weight matrix, FastKron uses $O(m+n)$ sequential steps while reducing the total work from $O(m^2n^2)$ to $O(mn(m+n))$. Thus, it matches the cubic scaling of GPTQ while exploiting richer curvature information. Moreover, FastKron is modular with respect to both the base quantizer and the Hessian estimator. We also provide practical benchmarks, consider a range of Hessian approximations that FastKron can be used with, and provide an efficient technique to compute these Hessians.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑