扩散Transformer的重要性感知低秩蒸馏
Importance-Aware Low-Rank Distillation of Diffusion Transformers
浏览论文内容
中文总结 AI 辅助
针对DiTs部署效率问题,提出SVDtrunc两步块级压缩方案,在40%-90%参数压缩下优于基准方法,68%参数预算仍近全性能,57%仍具竞争力,可与步骤蒸馏互补。
中文摘要 AI 辅助
扩散Transformer(DiTs)已成为高质量文本到图像生成的主流架构,但其规模给高效部署带来挑战。虽然截断奇异值分解(SVD)是参数缩减的合理工具,但大语言模型(LLMs)的研究表明,朴素低秩近似会引发灾难性失效。相反,我们发现DiTs中的截断SVD即使在大幅全局压缩下也会产生平滑的性能下降,冗余分布在整个网络的投影矩阵中,而非集中在少数Transformer块。基于这些见解,我们提出SVDtrunc,一种两步块级压缩方案:首先在全局参数预算下分配各块的秩,并通过截断SVD压缩重要性最低的块;随后通过模块化知识蒸馏和修正流目标对所有块进行微调。我们将SVDtrunc应用于[项目链接],压缩比例范围为原始参数数量的40%-90%。在GenEval、HPSv2和DPG三个基准测试中,我们的方法优于所有竞争方案。值得注意的是,与 prior 工作不同,我们在68%的参数预算下仍保持接近完整的性能,甚至在57%的参数预算下仍具竞争力。此外,我们表明SVDtrunc与步骤蒸馏互补,即使不进行微调也能取得良好结果,使其成为大型生成模型除扩散步骤缩减外效率提升的实用延续。项目页面:[项目链接]
英文摘要
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While truncated singular value decomposition (SVD) is a principled tool for parameter reduction, evidence from large language models (LLMs) suggests that naive low-rank approximation can cause catastrophic failure. In contrast, we find that truncated SVD in DiTs produces smooth degradation even under substantial global compression, with redundancy distributed across projection matrices throughout the whole network rather than concentrated in a few transformer blocks. Building on these insights, we introduce SVDtrunc, a two-step block-level compression scheme, first allocating ranks across blocks and compressing the least important ones via truncated SVD under a global parameter budget, and then fine-tuning all blocks with modular knowledge distillation and a rectified-flow objective. We apply SVDtrunc to FLUX.dev across compression levels ranging from 40-90% of the original parameter count. Across three benchmarks, GenEval, HPSv2, and DPG, we outperform all competing approaches. Notably, and in contrast to prior work, we retain near-full performance at 68% and remain competitive even at 57% of the original parameter budget. Furthermore, we show that SVDtrunc complements step distillation and achieves strong results even without fine-tuning, positioning it as a practical continuation of efficiency improvements beyond diffusion step reduction for large-scale generative models. Project page: https://vislearn.github.io/SVDtrunc/
发表机构
- Heidelberg University(海德堡大学)
- Zuse School ELIZA(楚思ELIZA学院)
- TU Darmstadt(达姆施塔特工业大学)
机构由 AI 辅助整理,请以论文原文为准。