混合迭代生成模型的训练后量化
Post-training Quantization for Hybrid Iterative Generative Models
浏览论文内容
中文总结 AI 辅助
针对混合迭代生成模型训练后量化易引发模型崩溃的问题,提出HyGenQ框架,通过分层聚类解耦与缩放重校准,成功将其量化至8位精度且性能优于现有方法。
中文摘要 AI 辅助
迭代生成模型(IGMs)涵盖自回归与扩散范式,结合二者的混合变体可实现出色的图像生成保真度。然而,其迭代推理会产生大量计算开销,使得训练后量化(PTQ)成为加速的可行方案,但直接将普通PTQ应用于混合IGMs会引发模型崩溃。通过分析这些失败情况,我们确定了两个关键挑战:激活值中的过多异常值(EOs)在保留正常精度与覆盖EOs之间形成无法调和的权衡,导致生成质量严重下降;微小量化误差引发的放大异常(AAs)不可预测地出现,造成校准与推理之间的不匹配,从而迭代触发模型崩溃。为应对这些挑战,我们提出HyGenQ,这是一个适用于混合IGMs的PTQ框架。HyGenQ包含分层聚类解耦(HCD)与缩放重校准(SR)两个模块:HCD通过多阶段聚类过程识别并解耦异常值通道,在保留正常值精度的同时有效隔离EOs,从而缓解性能下降;SR对超出高斯边界的AAs进行缩放,避免因激进截断引发的模型崩溃。大量实验表明,HyGenQ成功将代表性混合IGMs量化至8位精度(W8A8),显著优于现有基线方法,并验证了其在不同模型家族间的鲁棒性。
英文摘要
Iterative Generative Models (IGMs) span autoregressive and diffusion paradigms, and hybrid variants that couple them can achieve remarkable image-generation fidelity. However, their iterative inference incurs substantial computational overhead, making Post-training Quantization (PTQ) appealing for acceleration, while directly applying vanilla PTQ to hybrid IGMs can trigger model collapse. By analyzing these failures, we identify two critical challenges: Excessive Outliers (EOs) in the activations create an irreconcilable trade-off between preserving normal precision and covering EOs, resulting in severe degradation in generation quality; Amplified Anomalies (AAs) arising unpredictably from minor quantization errors, create a mismatch between calibration and inference, thus iteratively triggering model collapse. To address these challenges, we introduce HyGenQ, a PTQ framework for hybrid IGMs. HyGenQ comprises Hierarchical Cluster Decoupling (HCD) and Scaling Recalibration (SR). HCD identifies and decouples outlier channels via a multi-stage clustering process, effectively isolating EOs while maintaining normal value precision, thereby alleviating performance degradation. SR scales AAs beyond Gaussian Bound, thereby avoiding model collapse caused by aggressive truncation. Extensive experiments demonstrate that HyGenQ successfully quantizes representative hybrid IGMs to 8-bit precision (W8A8), significantly outperforming existing baselines and validating its robustness across different model families.
发表机构
- Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学研究所)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。