TASQ:用于扩散模型的时间自适应比特稀疏化量化
TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
浏览论文内容
中文总结 AI 辅助
针对扩散模型静态量化的计算开销问题,提出TASQ方法,结合时间精度引擎,在保持质量的同时显著降低执行周期。
中文摘要 AI 辅助
静态量化为每个去噪步骤分配一个权重精度,为了保持质量,该精度必须适配对量化最敏感的步骤,尽管许多其他步骤可以容忍更少的比特,得到的模型可能满足内存预算,但在整个去噪轨迹中反复承担最坏情况的算术开销。我们提出时间自适应比特稀疏化量化(TASQ)来分离这两种成本,TASQ存储一个共享的最高精度权重缓冲区,并学习一个时空最低有效位掩码,通过截断最低有效位为每个层和去噪阶段选择较低的有效精度,因此存储量由最坏情况保持固定,而比特运算(BitOPs)在敏感性较低的阶段减少,无需每个阶段的权重副本或运行时搜索。时间精度引擎将学习到的调度映射到比特串行执行,其中周期随有效精度缩放,且切换精度无测量到的周期开销。在PixArt-Sigma、SANA-1.6B和SDXL-Turbo上,TASQ实现了与静态量化相当的质量,同时计算量更小,结合时间精度引擎,它比静态量化减少25%至50%的执行周期,比朴素的静态8比特比特串行执行减少6.1至7.5倍,代码可在指定URL获取。
英文摘要
Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sensitive step, even though many other steps can tolerate fewer bits. The resulting model may satisfy its memory budget, but it repeatedly pays worst-case arithmetic throughout the denoising trajectory. We introduce Temporal-Adaptive Bit Sparsification Quantization (TASQ) to separate these two costs. TASQ stores one shared maximum-precision weight buffer and learns a Temporal-Spatial LSB Mask that selects a lower effective precision for each layer and denoising stage by truncating least-significant bits. Storage therefore remains fixed by the worst case, while BitOPs decrease at less sensitive stages without per-stage weight copies or runtime search. A Temporal-Precision Engine maps the learned schedule to bit-serial execution, where cycles scale with effective precision and switching precision has no measured cycle overhead. On PixArt-Sigma, SANA-1.6B, and SDXL-Turbo, TASQ achieves quality comparable to static quantization with less computation. Together with the Temporal-Precision Engine, it reduces execution cycles by 25 to 50 percent over static quantization and by 6.1 to 7.5x over a naive static 8-bit bit-serial execution. Code is available at https://github.com/seokho-han/tasq.