arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34732cs.CV

视觉生成器需要教师提供什么:重新思考表示对齐

What Visual Generators Need from Teachers: Rethinking Representation Alignment

Yongcong Wang, Hingchin Chen, Mingyu Fan, Shuo Jiang, Teer Zhang, Yucong Sun, Zijia Wang, Yiming Lu, Chengchao Shen

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出RARE方法,通过可恢复性差距选择教师层并动态加权token,加速扩散变换器训练,在ImageNet上以更少计算取得优于现有对齐基线的FID。

中文摘要 AI 辅助

表示对齐通过将模型(学生)的中间块拉向冻结预训练编码器(教师)的特征来加速扩散变换器训练。对齐哪个教师层以及对齐多长时间,仍然由惯例设定,而每种替代方案都需要一次训练运行的成本。我们发现,对齐在学生无法线性恢复教师特征的地方有帮助,而在学生已经与教师特征相似的地方则无益。由于深层教师层在很大程度上可从其下一层预测,我们隔离了每一层所增加的内容(其增量),并衡量未对齐的学生能恢复多少增量。学生从下往上填充教师的层级结构,并在接近顶部时停滞,我们称之为层级填充:即使在400K步之后,它也几乎无法恢复最深层的增量。可恢复性差距是增量中未恢复的部分,可从一次未对齐的检查点读取。在每次对齐一个教师层于一个块的短运行中,差距几乎重现了它们按FID改进的排名,而CKA(一种特征相似性度量)则大体上逆转了该排名。表示对齐与可恢复性估计(RARE)在训练前选择具有最大差距的教师层。在训练期间,它跟踪每个token到该层的剩余距离(差距的在线对应物),据此对token加权,并在平均距离停止下降时逐步淘汰损失。使用SiT-B/2在ImageNet 256×256上,RARE在无引导的情况下达到18.02的FID,有引导时达到4.46,领先于包括REPA、iREPA和HASTE在内的七个对齐基线。它也比iREPA少用14%的GPU小时进行训练。其FID在模型规模、教师、数据集和骨干网络方面均保持在iREPA之下。

英文摘要

Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet $256\times256$, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.

发表机构

  • Central South University(中南大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Tsinghua University(清华大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • SenseTime Research(商汤科技研究院)
  • Shandong University(山东大学)
  • Imperial College London(伦敦帝国学院)
  • University of Oxford(牛津大学)
  • Dell Technologies(戴尔科技)
  • University of International Relations(国际关系学院)

机构由 AI 辅助整理,请以论文原文为准。

↑