arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭示扩散Transformer中AdaLN-Zero的秘密

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

Jie Zhu, Mingyu Ding, Boqiang Duan, Leye Wang, Jingdong Wang

arXiv 2608.09438首次发表:更新:

发表机构

Peking University; UC Berkeley; Baidu(北京大学; 加州大学伯克利分校; 百度)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究扩散Transformer(DiT)中AdaLN-Zero性能优于AdaLN的原因,发现零初始化是关键要素,提出AdaLN-Gaussian初始化策略与SE-adaLN-Zero机制,经多数据集实验验证其有效性与泛化性。

AI 中文摘要

扩散Transformer(DiT)是一种快速兴起的图像生成架构,已获得大量关注。然而,尽管人们持续努力提升其性能,对DiT的理解仍较为浅显。本研究深入探究DiT内的关键条件机制AdaLN-Zero,其性能优于AdaLN。本研究分析驱动该性能的三个潜在要素:类SE结构、零初始化及“渐进式”更新顺序,其中零初始化被证实影响最大。基于此理解,我们提出一种经分析指导的初始化策略,名为AdaLN-Gaussian,它既用于实证验证分析结果,也作为一种实用初始化方法,可持续提升优化效率。另一方面,受类SE结构启发,我们引入一种改进的条件机制,称为SE-adaLN-Zero。在DiT框架下针对四个数据集开展的大量实验,尤其是在ImageNet1K上的实验,证实了AdaLN-Gaussian与SE-adaLN-Zero的有效性及泛化性。除类别到图像的生成外,我们还评估了这两种改进方法在文本到图像生成上的泛化能力。

英文摘要

Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improve its performance, the understanding of DiT remains superficial. In this work, we delve into and investigate a critical conditioning mechanism within DiT, adaLN-Zero, which achieves superior performance compared to adaLN. Our work studies three potential elements driving this performance, including an SE-like structure, zero-initialization, and a "gradual" update order, among which zero-initialization is proved to be the most influential. Building on this understanding, we propose an analysis-guided initialization strategy, termed adaLN-Gaussian, which serves both as an empirical validation of our analysis and as a practical initialization method that consistently improves optimization efficiency. On the other hand, inspired by the SE-like structure, we introduce an improved conditioning mechanism called SE-adaLN-Zero. Extensive experiments following DiT on four datasets, especially on ImageNet1K demonstrate the effectiveness and generalization of adaLN-Gaussian and SE-adaLN-Zero. Beyond class-to-image generation, we also evaluate the generalization of the two improved methods on text-to-image generation.

CommentsAccept by IEEE TPAMI 2026, camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑