发表机构
Carnegie Mellon University; Rice University(卡内基梅隆大学; 莱斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出跨视角类令牌对齐方法,在扩散 Transformer 中联合优化表征与生成,显著提升 ImageNet 线性探测和分割性能,同时保持或改善生成质量。
AI 中文摘要
生成式学习与表征学习之间的联系仍不对称:语义表征被用于改进扩散生成,而模型自身的表征往往被视为合成的副产品。我们探讨扩散模型能否在不牺牲生成质量的前提下,被训练以学习显著更强的语义表征。SelfFlow 通过将自监督补丁对齐引入流匹配朝此方向迈出了一步,但其主要收益仍在于更快的收敛和更好的生成。受 DINO 和 iBOT 启发,我们扩展了这一框架,加入跨视角类令牌对齐以进一步增强语义表征。具体而言,我们对每张图像形成两个独立加噪的双时间步观测,并将每个学生类令牌表征与来自另一观测的停止梯度 EMA 教师目标对齐。该目标与继承的流匹配和局部补丁目标联合优化。值得注意的是,尽管附加目标仅作用于类令牌,它同时增强了类令牌和补丁表征。与匹配的双视角基线相比,ImageNet 线性探测准确率在使用类令牌时提升 9.4%,使用均值池化补丁令牌时提升 10.1%,而冻结骨干的 VOC2012 分割提升 3.6 mIoU。这些表征增益在保持相当的 ImageNet 生成 FID 的同时实现。在文本到图像训练中,相同目标也改善了生成 FID,在匹配检查点处从 2.52 降至 2.37。我们的结果表明,表征不必仍是生成的副产品或仅作为改进生成的工具:它可以在扩散预训练中作为与生成并列的一等能力被直接优化。
英文摘要
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4\% using the class token and 10.1\% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.