结构化残差连接对扩散变换器的重要性
Structured Residual Connectivity Matters for Diffusion Transformers
AI总结:
本文提出扩散变换器的结构化残差连接,将被动求和转为主动检索机制,实现更快收敛和更优FID,仅需极少额外参数,显著提升生成质量。
AI中文摘要:
扩散变换器(DiTs)已确立为高保真图像合成的可扩展骨干网络。然而,与依赖刚性、手工设计的跳跃连接的基于U-Net的扩散模型不同,DiTs主要使用统一的残差流,将所有前层集成为单一状态。在这项工作中,我们重新思考扩散变换器中的残差连接,并提出将其从被动求和转变为针对图像去噪优化的主动检索机制。首先,我们对DiT的内部表示进行了系统分析,揭示了对早期层特征重用和对称层指导的潜在偏好。受此启发,我们引入了一种结构化连接设计,明确地将局部残差连接与长距离路径相结合。与静态跳跃连接或密集的全层路由不同,我们的方法使每个变换器块能够选择性地“关注”关键的早期表示,通过直接、可微分的跨深度路径动态检索空间和语义线索。实验表明,我们的自适应连接导致更快的收敛,训练迭代次数减少高达1.73倍,并且在额外参数少于0.1%的情况下,FID和视觉质量显著提升,进一步将强大的REPA-XL/2模型在无引导情况下从5.9 FID改进到4.34 FID,并在使用无分类器引导时达到1.39 FID。我们的研究结果表明,自适应跨层连接是扩散变换器中一个关键但未被充分探索的因素,并且纳入结构化信息路径为改进可扩展生成模型提供了一个简单而有效的方向。
英文摘要:
Diffusion Transformers (DiTs) have established themselves as a scalable backbone for high-fidelity image synthesis. However, unlike U-Net based diffusion models that rely on rigid, hand-crafted skip connections, DiTs predominantly use a uniform residual stream that integrates all preceding layers as a monolithic state. In this work, we rethink residual connections in diffusion transformers and propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising. First, we conduct a systematic analysis of DiT's internal representation, revealing a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by this, we introduce a structured connectivity design that explicitly integrates local residual connections with long-range pathways. Instead of static skip connections or dense all-layer routing, our method enables each transformer block to selectively ``attend'' to critical earlier representations, dynamically retrieving spatial and semantic cues through direct, differentiable cross-depth paths. Experiments show that our adaptive connectivity leads to faster convergence, with up to $1.73\times$ fewer training iterations, and significant gains in FID and visual quality with less than $0.1\%$ additional parameters, further improving a strong REPA-XL/2 model from $5.9$ to $4.34$ FID without guidance and reaching $1.39$ FID with classifier-free guidance. Our findings suggest that adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that incorporating structured information pathways provides a simple and effective direction for improving scalable generative models.