arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环扩散Transformer

Looped Diffusion Transformer

Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang

arXiv 2609.40305首次发表:更新:

发表机构

SenseTime Research; LeapLab, Tsinghua University(商汤科技研究院; 清华大学LeapLab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出循环扩散Transformer(Looped-DiT),通过共享块循环计算与深度监督及自调制注意力,在固定参数下超越更大模型,实现高效迭代生成。

AI 中文摘要

传统上,改进文本到图像模型依赖于增加模型规模或去噪步骤的数量。在这项工作中,我们探索了一种替代性的计算扩展方式,即在每个去噪步骤内重复运行共享的Transformer块,从而在保持参数数量不变的情况下有效增加计算深度。这种循环计算使得内部表示能够在不使用显式推理令牌的情况下进行迭代细化。然而,简单的循环未能持续改善图像质量。我们将此问题追溯到中间循环之间的弱监督以及逐渐侵蚀局部信息的无约束注意力更新。为了克服这些挑战,我们提出了循环扩散Transformer(Looped-DiT),它结合了中间循环的深度监督和自调制注意力,以稳定循环特征更新。在参数匹配和计算匹配的设置下,Looped-DiT始终优于非循环基线。值得注意的是,一个260M参数的循环模型可以在多个文本到图像基准上超越大6.5倍的模型,同时推理计算需求降低4.9倍。除了这一性能提升,我们还发现循环计算可以为扩散模型提供一种更有效的迭代计算形式,在固定推理预算下,增加循环深度比增加去噪步骤带来更大的收益。此外,更深的循环可以逐步纠正早期循环中的错误,表现出暗示潜在推理的行为。总之,这些结果表明循环计算为扩展视觉生成模型提供了一种有前景的方式。

英文摘要

Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. This looped computation enables iterative refinement of internal representations without explicit reasoning tokens. However, naive looping fails to consistently improve image quality. We trace this problem to weak supervision across intermediate loops and unregulated attention updates that progressively erode local information. To overcome these challenges, we propose Looped Diffusion Transformer (Looped-DiT), which combines deep supervision across intermediate loops with self-modulating attention to stabilize looped feature updates. Under matched-parameter and matched-compute settings, Looped-DiT consistently outperforms non-looped baselines. Notably, a 260M-parameter looped model can surpass a model 6.5x larger across multiple text-to-image benchmarks while requiring 4.9x lower inference compute. Beyond this performance gain, we find that looped computation can offer a more effective form of iterative computation for diffusion models, with increasing loop depth yielding larger gains than adding more denoising steps under a fixed inference budget. Furthermore, deeper loops can progressively correct mistakes made in earlier loops, exhibiting behaviors suggestive of latent reasoning. Together, these results show that looped computation offers a promising way to scale visual generation models.

Comments21 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑