arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33496cs.LGcs.AIcs.CV

Chameleon:用于高效扩散模型的动态格式适配器

Chameleon: Dynamic Format Adapter for Efficient Diffusion

Arnab Sanyal, Sandeep Chinchali

首次发表
浏览论文内容

中文总结 AI 辅助

针对扩散模型训练后量化中固定数字格式导致性能不佳的问题,提出Chameleon框架,将格式作为离散变量按通道和时间步动态选择,在多种模型和位宽下取得最佳FID。

中文摘要 AI 辅助

训练后量化(PTQ)是在内存受限的加速器上运行现代扩散模型的标准方式,然而现有的扩散模型PTQ方案都预先固定数字格式,仅调整缩放因子、零点或逐层位宽。在固定位宽下,最佳格式取决于所编码的分布,而该分布在权重通道之间、层之间以及扩散时间步之间均存在差异,其中激活分布从重尾且噪声主导的状态滑动到紧密聚集且结构化的状态。我们提出Chameleon,一个PTQ框架,它保持位宽固定,将格式本身视为离散变量,按权重通道和按(层,时间步桶)激活张量进行选择。激活格式来自{INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8},通过两个廉价统计量(经验峰度和闭式扩散信噪比)提前选择并存储在查找表中;权重格式在8位下来自{INT8, MXINT8},在4位下来自{INT4, NF4, FP4 E2M1, MXINT4, MXFP4},通过重构误差离线选择。一种架构分支将相同的选择层适配到多步UNet、单步蒸馏模型和扩散Transformer。在COCO-2014上的SDXL、SDXL-Turbo和PixArt-α上,Chameleon在所有六种骨干网络×位宽设置中实现了最佳FID,CLIP分数与FP16参考值相差在0.24以内,并且在W4A8设置下是所有量化方法中最好的。

英文摘要

Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best format depends on the distribution being encoded, and that distribution differs across weight channels, across layers, and along the diffusion timestep, where activation distributions slide from heavy-tailed and noise-dominated to tightly clustered and structured. We propose Chameleon, a PTQ framework that holds the bit-width fixed and treats the format itself as a discrete variable, chosen per weight channel and per (layer, timestep bucket) activation tensor. Activation formats come from {INT8, FP8 E4M3, FP8 E5M2, MXFP8, MXINT8}, selected ahead of time from two cheap statistics (empirical kurtosis and the closed-form diffusion SNR) and stored in a lookup table; weight formats come from {INT8, MXINT8} at 8 bits or {INT4, NF4, FP4 E2M1, MXINT4, MXFP4} at 4 bits, selected offline by reconstruction error. An architectural fork adapts the same selection layer to multi-step UNets, single-step distilled models, and Diffusion Transformers. Across SDXL, SDXL-Turbo, and PixArt-$α$ on COCO-2014, Chameleon achieves the best FID in all six backbone $\times$ bit-width settings, with CLIP within 0.24 of the FP16 reference and the best of all quantized methods at $W_{4}A_{8}$.

发表机构

  • UT SWARM Lab(得克萨斯大学西南偏南实验室)
  • The University of Texas, Austin(得克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑