嵌入预测有助于图像生成
Embedding Prediction Helps Image Generation
- University of Michigan(密歇根大学)
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出NEPA方法,通过预测下一嵌入作为扩散变换器的条件,结合多嵌入预测与REPA,在ImageNet上以约三分之一的训练计算达到FID 1.32。
AI中文摘要:
在扩散变换器中,类别标签或文本提示被嵌入一次,并且在每个去噪步骤中重复使用相同的条件。我们提出疑问:预测的嵌入能否作为此条件使用。下一嵌入预测自回归(NEPA)训练一个变换器来预测序列中的下一个连续嵌入。在生成过程中,干净图像跟随噪声图像,因此其嵌入是条件和噪声图像之后的下一个嵌入。我们训练一个NEPA模型,通过多嵌入预测一次性预测它们,并在嵌入条件生成中,一个DiT生成器以这些预测为条件,在每个去噪步骤中重新计算,因此条件信号适应当前的噪声状态。在类别条件的ImageNet $256\ imes256$上的实验研究了生成器的条件、多嵌入预测的设计以及两个模型的扩展。NEPA模型在每个采样步骤中增加了一个第二个网络;结合REPA,我们的最终模型NEPA-DiT-XL达到了1.32的FID,使用的训练计算量约为REPA的三分之一。
英文摘要:
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.