arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

扩散变换器的训练后剪枝

Post-Training Pruning for Diffusion Transformers

Chengzhi Hu, Xuewen Liu, Jing Zhang, Mengjuan Chen, Zhikai Li, Qingyi Gu

arXiv 2607.00927首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散变换器(DiTs)计算开销大的问题,提出DiT-Pruning方法,通过能量感知的显著性度量与聚类感知的剪枝粒度,在50%稀疏度下仅损失0.001 CLIP分数。

AI 中文摘要

扩散变换器(DiTs)在图像生成中表现出色,但存在显著的计算开销和资源消耗。训练后剪枝提供了一种有前景的解决方案;然而,由于DiTs独特的架构设计和参数分布,传统剪枝方法不适用,导致性能显著下降。具体而言,先前为LLMs开发的方法通过一系列近似推导度量,放大了权重在显著性度量中的相对贡献。此外,DiTs中的权重幅度显著大于LLMs中的权重。而且,现有剪枝粒度忽略了模型结构的变化。在本文中,我们提出DiT-Pruning,通过引入定制化的显著性准则和剪枝粒度来提升剪枝性能。我们设计了一种新颖的度量,从能量角度平衡权重和激活的贡献,从而更有效地识别重要元素。此外,我们观察到二维权重空间中的独特聚类模式。据此,我们采用聚类感知的剪枝粒度,实现有效的稀疏分配。在各种DiTs上的广泛评估表明,我们的方法一致地保持了图像质量,特别是在高稀疏度下。对于MJHQ上512x512分辨率的FLUX.1-dev,DiT-Pruning在50%稀疏度下仅损失0.001的CLIP分数,显著优于最近的剪枝方法。

英文摘要

Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhead and resource consumption. Post-training pruning offers a promising solution; however, due to DiTs' unique architectural design and parameter distribution, traditional pruning methods are inapplicable, leading to significant performance degradation. Specifically, prior methods developed for LLMs, which derive metrics through a series of approximations, amplify the relative contribution of weights in the saliency metric. In addition, weights in DiTs exhibit significantly larger magnitudes than those in LLMs. Moreover, existing pruning granularity overlooks variations in model structures. In this paper, we propose DiT-Pruning, which improves pruning performance by introducing customized saliency criteria and pruning granularity. We design a novel metric that balances the contributions of weights and activations from an energy-based perspective, enabling more effective identification of important elements. Furthermore, we observe distinct clustering patterns in the two-dimensional weight space. Accordingly, we adopt a clustering-aware pruning granularity, enabling effective sparse allocation. Extensive evaluations on various DiTs show that our method consistently preserves image quality, especially under high sparsity. For FLUX.1-dev at 512x512 resolution on MJHQ, DiT-Pruning achieves only a 0.001 loss in CLIP score at 50% sparsity, dramatically outperforming recent pruning methods.

Comments15 pages, 13 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑