arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01298cs.CVcs.AIcs.LG

UDT:通过数据自适应令牌缩减协调U-Net与扩散Transformer

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

Junno Yun, Yaşar Utku Alçalar, Mehmet Akçakaya

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出UDT架构,通过数据自适应令牌合并协调U-Net与DiTs,在ImageNet图像生成任务中收敛速度更快、FID指标更优,为DiTs提供了新的高效骨干网络。

中文摘要 AI 辅助

扩散Transformer(DiTs)凭借其可扩展性和对多模态任务的适应性,已成为生成建模领域的核心架构。DiTs由各向同性Transformer块组成,会在网络深度方向逐步学习表征,其中去噪目标会驱动后续层专注于精细细节的重建,这导致表征质量下降,且编码器-解码器行为失衡。现有方法如表征对齐(REPA)通过训练正则化来强化早期表征以缓解该问题;而U-Net风格的DiT架构则引入显式多尺度编码器-解码器结构以提升收敛性,但这类架构基于标准U-Net的空间下采样可学习算子构建,该算子与Transformer架构适配性不佳,会引入效率问题并与交叉注意力、表征正则化等组件产生兼容性问题。本研究提出UDT,一种U-Net扩散Transformer,它结合了DiTs的表征能力与U-Nets的编码-解码优势,通过数据自适应令牌合并实现下采样与上采样,同时保留DiT的令牌维度。我们的基线UDT架构在所有模型规模下均优于现有U-Net DiTs,且性能与REPA相当。此外,在256×256 ImageNet数据集上,XL规模模型的UDT在40个训练周期内(约40倍更快的收敛速度),便在无分类器引导(w/o CFG)场景下超过了SiT在1400个周期达到的7.9 FID指标。最后,该模型结合分类器引导(CFG)实现了出色的图像生成性能,使用SD-VAE时在320个周期达到FID 1.38,使用VA-VAE时在500个周期达到FID 1.35,为DiTs提供了具有强实证优势的新骨干网络。

英文摘要

Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder-decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder-decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose UDT, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding-decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs (~ 40x faster convergence) for XL model size on 256x256 ImageNet. Finally, it achieves strong image generation performance with CFG, reaching FID of 1.38 (320 epochs) with SD-VAE and 1.35 (500 epochs) with VA-VAE, providing a new backbone for DiTs with strong empirical benefits.

发表机构

  • University of Minnesota(明尼苏达大学)

机构由 AI 辅助整理,请以论文原文为准。

↑