AI 中文总结
该研究针对扩散Transformer(DiTs)表示学习机制不明的问题,提出加权多样性分数(WDS)指标,进而构建DiverseDiT++框架提升表示多样性,在ImageNet数据集上实现了性能提升与收敛加速。
AI 中文摘要
扩散Transformer(DiTs)的最新进展凭借其出色的可扩展性,在视觉合成领域取得了显著进步。为提升DiTs捕捉有意义内部表示的能力,近期如REPA这类工作会引入外部预训练编码器来进行表示对齐。然而,学界对DiTs内部表示学习的潜在机制仍知之甚少。为此,本文首先通过量化各模块表示的多样性,对DiTs的表示动态展开系统性分析。具体而言,我们提出了一种名为加权多样性分数(WDS)的新指标,用于衡量不同模块间的表示差异。通过在多种设置下对内部表示的演化及影响展开广泛研究,我们发现跨模块的表示多样性是DiTs实现有效表示学习的关键因素。更重要的是,WDS在各类设置、模型规模和训练阶段下均与合成质量呈现强相关性(与log(FID)的皮尔逊相关系数r=-0.869),表明其可作为反映模型性能的指标,也可作为模型优化的原则性指导。基于这一关键发现,我们提出了DiverseDiT++这一新型框架,该框架明确推动多样化表示学习。具体来说,我们的方法引入了长残差连接以实现各模块间输入表示的多样化,并加入表示多样性损失以鼓励各模块学习不同的特征。在ImageNet 256×256和512×512数据集上开展的大量实验表明,将DiverseDiT++应用于不同规模的各类主干网络时,均能实现稳定的性能提升与收敛加速。
英文摘要
Recent advances in Diffusion Transformers (DiTs) have enabled remarkable progress in visual synthesis, benefiting from their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as REPA incorporate external pretrained encoders for representation alignment. However, the underlying mechanisms governing representation learning within DiTs remain poorly understood in the community. To this end, this paper first presents a systematic analysis of the representation dynamics of DiTs via quantifying the diversity of block-wise representations. Specifically, we introduce a novel metric, termed the Weighted Diversity Score (WDS), to measure the representational discrepancies across different blocks. Through extensive investigations on the evolution and influence of internal representations under various settings, we reveal that representation diversity across blocks is a critical factor for effective representation learning in DiTs. More importantly, WDS exhibits a strong correlation with synthesis quality across diverse settings, model scales, and training stages (Pearson's $r=-0.869$ with $\log(\text{FID})$), suggesting its potential as an indicator to reflect model performance and a principled guide for model optimization. Based on this key finding, we propose DiverseDiT++, a novel framework that explicitly promotes diverse representation learning. Concretely, our method incorporates long residual connections to diversify input representations across blocks and a representation diversity loss to encourage blocks to learn distinct features. Extensive experiments on ImageNet $256\times256$ and $512\times512$ demonstrate that our DiverseDiT++ yields consistent performance gains and convergence acceleration when applied to different backbones with various sizes,...
Comments35 pages, 32 figures