AI 中文总结
针对学习型图像编码均匀INT8量化性能受限问题,提出HaTQ方法,通过哈达玛重参数化适配INT8量化,实验显示其在不同设置下性能提升且推理效率高。
AI 中文摘要
均匀INT8量化适用于部署学习型图像编码(LIC),但由于重尾张量和较大的通道间差异,其率-失真(R-D)性能常受限制。现有方法主要通过混合精度或非均匀码本调整量化器。我们提出哈达玛变换域量化(HaTQ),该方法在量化前采用正交哈达玛重参数化,将原始域中的权重和激活响应跨通道重新分配。该重参数化保留了每个线性算子的原始函数映射,同时使其权重和激活更适配均匀INT8量化。HaTQ提供两种互补形式:双哈达玛量化对输入激活和权重均进行变换,仅权重哈达玛量化仅对权重进行变换。这种区分很重要,因为恒定哈达玛基可在敏感层中相干积累非零通道均值并扩大激活范围。我们通过离线分析识别这些敏感层,为每层分配合适的形式且不依赖输入分支。HaTQ支持训练后量化(PTQ)和量化感知训练(QAT),使用均匀INT8量化器,且与仅整数执行兼容。在代表性LIC架构和数据集上的实验表明,其在不同量化设置下均实现一致提升,所得QAT模型进一步优于竞争的混合精度和非均匀量化方法,TensorRT部署结果显示出实用的INT8推理效率,源代码将公开发布。
英文摘要
Uniform INT8 quantization is attractive for deploying learned image coding (LIC), but its rate--distortion (R--D) performance is often limited by heavy-tailed tensors and large inter-channel variations. Existing methods mainly adapt the quantizer through mixed precision or non-uniform codebooks. We propose Hadamard-Transform-domain Quantization (HaTQ), which uses orthogonal Hadamard reparameterization before quantization to redistribute weight and activation responses in the original domain across channels. The reparameterization preserves the original function mapping of each linear operator, while making its weights and activations more amenable to uniform INT8 quantization. HaTQ provides two complementary forms. Double-Hadamard quantization transforms both the input activations and weights, whereas weight-only Hadamard quantization transforms only the weights. This distinction is important because the constant Hadamard basis can coherently accumulate a nonzero channel mean and enlarge the activation range in sensitive layers. We identify these sensitive layers through offline profiling and assign the appropriate form to each layer without input-dependent branching. HaTQ supports both post-training quantization (PTQ) and quantization-aware training (QAT), uses uniform INT8 quantizers, and is compatible with integer-only execution. Experiments on representative LIC architectures and datasets demonstrate consistent improvements across different quantization settings. The resulting QAT models further outperform competing mixed-precision and non-uniform quantization methods. TensorRT deployment results demonstrate practical INT8 inference efficiency. The source code will be publicly released.