AI 中文总结
针对去噪扩散Transformer训练收敛慢的问题,提出无结构参数亲和力正则化SPARE,利用跨图像标记对的结构实现正则化,在ImageNet数据集上取得优异生成性能,内存开销极小。
AI 中文摘要
去噪扩散Transformer实现了出色的生成质量,但训练过程收敛缓慢。对其内部表示进行正则化已成为一种有效的加速手段,不过现有方法分为两类,成本各有优劣:基于目标的方法通过将表示与外部特征对齐来增强表示,这需要外部编码器和可学习的投影头来桥接特征空间;无目标方法则完全不依赖参考,只能在样本或层之间排斥模型自身的特征,丢弃了数据包含的所有结构。先前研究表明,驱动对齐增益的是空间结构而非全局语义,因此我们探究这种结构是否可直接作为目标,且是否不仅存在于单张图像内,还存在于不同图像之间。我们的核心见解是,干净数据的隐变量已在其标记之间的关系中携带了这种结构,其中关系是两个标记之间的相似度,是无需投影头即可在特征空间间比较的单个标量。我们提出无结构参数亲和力正则化(SPARE),这是一种将中间标记的成对亲和力与干净隐变量的成对亲和力匹配的正则化器。为充分利用这种结构,SPARE将匹配扩展到跨图像的标记对,即先前无目标方法默认排斥的那些对,并通过单一学习目标校准这两种关系类型。在ImageNet $256 \times 256$ 数据集上,使用SiT骨干网络,在40万次迭代的匹配预算下,SPARE未添加编码器、头或参数,仅增加0.08GB训练内存,却在所有测试设置的无参数正则化器中取得最低FID,恢复了REPA 37%至54%的FID降低幅度,且与REPA结合时表现更优,在无分类器引导下100万次迭代时达到FID 1.90。
英文摘要
Denoising diffusion transformers achieve strong generation quality but converge slowly during training. Regularizing their internal representations has emerged as an effective accelerator, yet existing methods split into two families with complementary costs. Target-based methods strengthen representations by aligning them to external features, which requires an external encoder and a learnable projection head to bridge feature spaces. Target-free methods hold no reference at all, and can only repel the model's own features across samples or layers, discarding whatever structure the data contains. Prior work suggests that spatial structure, rather than global semantics, drives the gains of alignment. We therefore ask whether such structure can serve as a target directly, and whether it exists not only within an image but across images. Our key insight is that the clean data latent already carries this structure in the relations among its tokens, where a relation is the similarity between two tokens, a single scalar comparable across feature spaces without a projection head. We propose Structural Parameter-free Affinity Regularization (SPARE), a regularizer that matches the pairwise affinities of intermediate tokens to those of the clean latents. To exploit this structure fully, SPARE extends the matching to token pairs across images, precisely the pairs that prior target-free methods repel by default, and calibrates both relation types with a single learning objective. On ImageNet $256 \times 256$ with SiT backbones under matched 400K-iteration budgets, SPARE adds no encoder, head, or parameters and only 0.08 GB of training memory, yet attains the lowest FID among parameter-free regularizers in every tested setting, recovers 37 to 54\% of REPA's FID reduction, and improves over REPA when combined with it, reaching FID 1.90 under classifier-free guidance at 1M iterations.
CommentsPreprint