arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层误差归因用于快速且鲁棒的混合精度训练后量化

Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization

Samy Houache, Yann Traonmilin, Jean-François Aujol

arXiv 2610.09877首次发表:更新:

发表机构

Univ. Bordeaux; Bordeaux INP; Thales AVS; CNRS(波尔多大学; 波尔多国立理工学院; 泰雷兹AVS; 法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于分层概率误差分析的可分离评分方法,用于快速鲁棒的混合精度训练后量化,在去噪和扩散模型上匹配或超越现有基线,并显著提升对损坏校准的鲁棒性。

AI 中文摘要

混合精度训练后量化是一种网络压缩方法,在全局内存预算下使用小型校准集逐层分配比特。主要困难在于克服分配问题的组合性质,并管理对小型且可能被破坏的数据集的敏感性。因此,一种高效的分配方法应计算快速,并在校准数据被破坏时保持模型质量。为设计此类方法,我们对量化误差进行了分层概率分析,将传播误差与给定层引入的局部扰动分开。我们利用该局部项为简单分配算法构建可分离评分,无需外部求解器。我们方法的概率性质带来了对损坏数据的鲁棒性。在DRUNet的去噪任务中,平均每权重4比特预算下,我们的方法在干净校准下匹配或改进了最先进的混合精度基线,并且对损坏校准更鲁棒,在测试的损坏情况下PSNR增益高达7.5 dB。实验显示,与所研究的基线相比,比特分配加速了28倍至2570倍。对于量化扩散模型,我们的实验表明直接应用我们的框架也改进了最先进水平。

英文摘要

Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no external solver. The probabilistic nature of our approach brings robustness to corrupted data. On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions. Experiments show bit-allocation speed-ups from 28x to 2,570x over the studied baselines. For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑