面向量化矩阵乘法的乘积感知确定性舍入
Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
浏览论文内容
中文总结 AI 辅助
研究量化矩阵乘法中的确定性乘积感知舍入,提出零空间约简与条件期望补全算法,在动态激活和静态权重场景下降低乘积误差,实验验证优于舍入到最近。
中文摘要 AI 辅助
标量舍入决策通过矩阵乘法相互作用。我们研究在缩放因子、裁剪边界和网格固定后,每个活动标量在相邻层级之间选择的确定性乘积感知舍入。对于动态激活舍入,零空间约简在保留松弛乘积的同时,最多留下$r$个分数决策,其中$r$是活动间隙加权权重块的秩。条件期望补全给出一个确定性多项式时间算法,其平方乘积误差至多为$\text{OPT}_{\text{dyn}}+r\nu_{\text{max}}^2/4$,其中$\text{OPT}_{\text{dyn}}$是最佳可容许误差,$\nu_{\text{max}}$是该块的最大行范数。对于可复用的静态权重,精确的期望乘积损失度量是带有固定输出偏差的非中心输入二阶矩;自由偏差重新校准产生中心协方差。即使在秩为一的情况下,精确优化也是NP难的。在$K=1024$、$r=16$的平衡块中,条件期望补全达到抖动归一化中位误差$0.010$,而舍入到最近为$0.899$。裁剪感知初始化在百分之十裁剪下将中位归一化误差降低$43.4$倍。留出法Digits实验表明,在所有四种测试的位宽和校准大小设置下,保留输入均值或校正输出偏差相对于舍入到最近改善了中位乘积误差。
英文摘要
Scalar rounding decisions interact through matrix multiplication. We study deterministic product-aware rounding after scales, clipping bounds, and grids are fixed, with each active scalar choosing between adjacent levels. For dynamic activation rounding, null-space reduction preserves the relaxed product while leaving at most $r$ fractional decisions, where $r$ is the rank of the active gap-weighted weight block. Conditional-expectation completion gives a deterministic polynomial-time algorithm with squared product error at most $ \mathrm{OPT}_{\mathrm{dyn}}+rν_{\max}^2/4$, where $\mathrm{OPT}_{\mathrm{dyn}}$ is the best admissible error and $ν_{\max}$ is the largest row norm of that block. For reusable static weights, the exact expected product-loss metric is the uncentered input second moment with fixed output bias; free bias recalibration yields the centered covariance. Exact optimization is NP-hard even at rank one. In balanced blocks with $K=1024$ and $r=16$, conditional- expectation completion attains a dither-normalized median error of $0.010$, compared with $0.899$ for round-to-nearest. Clipping-aware initialization reduces median normalized error by a factor of $43.4$ at ten-percent clipping. Held-out Digits experiments show that retaining the input mean or correcting the output bias improves median product error over round-to-nearest in all four tested bit-width and calibration-size settings.