arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28270cs.CLcs.LG

预测量化价格:在部署前选择 PTQ 配置

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu

AI总结:

提出基于定价层输出误差的PTQ配置选择方法,利用下游曲率计算量化价格,统一比较不同格式、粒度和变换,实现部署前预算引导的配置选择。

AI中文摘要:

权值空间的后训练量化(PTQ)必须在完成量化模型揭示其输出分布漂移之前,选择有限格式、粒度、量化器族、变换和比特数。现有的 PTQ 方法预测了这种退化的关键部分,包括重建误差、Hessian 敏感性、变换效应和下游损失,但这些部分通常在固定量化几何或独立配置族内进行评分。我们将权值空间 PTQ 表述为使用定价的层输出误差进行部署前配置选择。每个可接受的层配置被视为具有部署成本的误差生成器,这产生了层输出误差协方差 $\boldsymbolΣ_l(α_l)$,而全精度模型通过下游曲率对该协方差定价,$\widehatρ_l(α_l)=\frac{1}{2}\operatorname{Tr}\left(\widehat{\mathbf{H}}_l\\,\widehat{\boldsymbolΣ}_l(α_l)\right)$。该价格源自全精度到量化前向 KL 散度,其一阶项在参考模型处抵消。它将重建和诊断分数转化为丢弃价格因子的简化代理,而有限格式、码本、粒度和等效变换通过它们引起的协方差和付出的成本成为可比较的候选。随后,迹缩减产生校准时间价格表和预算引导的价格导向选择器,使固定几何比特分配成为特例而非组织问题。

英文摘要:

Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization geometry or inside separate configuration families. We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error. Each admissible layer configuration is treated as an error generator with a deployment cost, which induces a layer-output error covariance $\boldsymbolΣ_l(α_l)$, and the full-precision model prices that covariance by downstream curvature, $\widehatρ_l(α_l)=\frac{1}{2}\operatorname{Tr}\left(\widehat{\mathbf{H}}_l\,\widehat{\boldsymbolΣ}_l(α_l)\right)$. The price follows from full-precision-to-quantized forward KL, whose first-order term cancels at the reference model. It turns reconstruction and diagonal scores into reduced proxies that drop price factors, while finite formats, codebooks, granularities, and equivalent transformations become comparable candidates through the covariances they induce and the costs they pay. A trace reduction then yields a calibration-time price table and a budgeted price-guided selector, making fixed-geometry bit allocation a special case rather than the organizing problem.

↑