发表机构
ENPC, IP Paris; Ecole Polytechnique, IP Paris; AMIAD; UC Berkeley(巴黎理工学院、巴黎IP大学; 巴黎综合理工学院、巴黎IP大学; AMIAD机构; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对流匹配模型训练的现有问题,提出直接预测DINO预训练表示的方法,消除了对第二条ODE及特定去噪调度的需求,在ImageNet上以更少epoch实现了更优生成质量。
AI 中文摘要
对DINO等预训练表示进行协同去噪可大幅提升流匹配模型的训练速度与质量,但这会引入第二条去噪轨迹,且需要精心设计的调度方案。我们提出了一种更简单的替代方案:直接预测预训练表示,再让模型以自身预测为条件。这消除了对第二条常微分方程(ODE)及任何特定表示去噪调度的需求,同时保留了表示引导的优势。我们的方法收敛速度显著更快,且在FID分数衡量下实现了更优的生成质量。在ImageNet上,它在隐空间中以比现有方法少2倍的epoch数超越了现有技术水平;在像素空间中,它将FID分数较可比现有方法提升了20%以上。这些结果支持了一个简单原则:不要去噪你可以预测的内容。我们的代码可在该httpsURL公开获取。
英文摘要
Co-denoising pretrained representations such as DINO can substantially improve the training speed and quality of flow matching models, but it introduces a second denoising trajectory and requires carefully designed schedules. We propose a simpler alternative: predict the pretrained representation directly, then condition the model on its own prediction. This removes the need for a second ODE and any representation-specific denoising schedules, while retaining the benefits of representation guidance. Our approach converges substantially faster and achieves better generation quality as measured by FID score. On ImageNet, it outperforms the state of the art in latent space at 2x fewer epochs than prior methods; in pixel space, it improves FID over comparable prior methods by more than 20%. These results support a simple principle: do not denoise what you can predict. Our code is openly available at https://github.com/arijit-hub/dino_forcing.