arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

量化矩阵乘法的收缩规范预处理

Contraction-Gauge Preconditioning for Quantized Matrix Multiplication

Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal

arXiv 2607.18745首次发表:更新:

发表机构

UT-Battelle, LLC; Oak Ridge National Laboratory(UT-巴特尔有限责任公司; 橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究两个因子均量化时\(C = AB\)的低精度计算,利用乘积保持等价关系制定收缩规范预处理,推导可计算选择统计量,给出上界用于排序候选者,并通过实验验证该方法在降低乘积误差等方面的有效性。

AI 中文摘要

我们研究了两个因子均量化时\(C = AB\)的低精度计算。在具有已知方差场的独立、零均值逐元素误差下,我们推导了期望平方乘积误差的精确有限维恒等式,该恒等式对非过载减法抖动和独立随机舍入精确成立,我们还通过实证评估了确定性四舍五入到最接近值(RTN)。利用乘积保持等价关系\(AB=(A^T)(T^{-1}B)\),我们制定了收缩规范预处理:在量化之前联合选择因子表示及其共享模式。预处理可减少乘积误差,但可能需要额外的相反操作数的变换、量化副本:共享变换需要一个副本,特定于块的变换每个块最多需要一个副本。在正对角规范(折叠)的有界族内,一个几何程序计算全局最优共享折叠,一个线性程序决定恒等折叠是否已经是最优的。对于其他族,我们推导了可计算的选择统计量——用于缩放的尾指数、用于划分的轮廓扩展、用于旋转的相干性和加权格拉姆能量、用于层次深度的切片能量协方差——并给出了用于对启发式候选者进行排序的上界。在一个经过训练的三块图像分类器的十二个线性乘积中,抖动模型预测与确定性RTN误差之间的乘积内秩中位数相关性在8位时为0.937,在4位时为0.918。几何程序计算的折叠在几何均值上比恒等折叠将保留的乘积误差降低了18.0%(8位)和20.5%(4位),在两种精度下以及在十二个乘积中的十个上击败了SmoothQuant风格的网格基线,并将组合对数似然均方误差降低了15.4%和26.4%。因此,我们提供了精确的随机乘积误差计算、对角族内的认证选择,以及一个用于评估RTN下可重用变换候选者的共同目标。

英文摘要

We study low-precision computation of C=AB with both factors quantized. We derive an exact finite-dimensional identity for the expected squared product error under independent, zero-mean entrywise errors with known variance fields; it holds exactly for non-overloading subtractive dither and for independent stochastic rounding, and we empirically assess deterministic round-to-nearest (RTN). Using the product-preserving equivalence AB=(AT)(T^{-1}B), we formulate contraction-gauge preconditioning: jointly choosing a factor representation and its sharing pattern before quantization. Preconditioning can reduce product error but may require extra transformed, quantized copies of the opposite operand: a shared transform needs one copy, a block-specific transform up to one per block. Within the bounded family of positive diagonal gauges (folds), a geometric program computes a globally optimal shared fold and a linear program decides whether the identity fold is already optimal. For other families we derive computable selection statistics -- tail index for scaling, profile spread for partitioning, coherence and weighted-Gram energy for rotations, slice-energy covariance for hierarchy depth -- with upper bounds for ranking heuristic candidates. Across twelve linear products from a trained three-block image classifier, median within-product rank correlations between dither-model predictions and deterministic-RTN errors are 0.937 at 8 bits and 0.918 at 4 bits. The GP fold cuts held-out product error over the identity fold by 18.0% (8-bit) and 20.5% (4-bit) in geometric mean, beats a SmoothQuant-style grid baseline at both precisions and on ten of twelve products, and lowers composed logit MSE by 15.4% and 26.4%. We thus provide exact stochastic product-error accounting, certified selection within the diagonal family, and a common objective for evaluating reusable transform candidates under RTN.

Comments50 pages, 13 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑