AI 中文总结
提出蒸馏后精炼的端到端潜在展开方法,用FiST架构在ImageNet-256上以少步生成达到FID 1.11,无需显式Fréchet损失。
AI 中文摘要
迭代生成在步骤间构成一个联合优化问题,因为中间预测会塑造后续计算,并最终决定最终输出分布。从预训练的扩散和流匹配模型中蒸馏出的少步生成器,使得这种优化在计算上可以端到端地进行。我们基于这一机会提出了一种“先蒸馏后精炼”的方法,利用教师模仿为潜在空间中的少步展开建立强初始化,然后转向针对真实数据对完整潜在展开进行端到端精炼。我们引入了FiST(Flow-in-Stage Transformer),一种架构,它使用共享Transformer在几个阶段中组合学习到的潜在状态转换,并可选地提供跨阶段隐藏通信。蒸馏阶段沿教师轨迹对选定状态应用教师强制回归;精炼阶段则用对抗性和辅助分类目标替代这种监督,作用于最终潜在输出。一个可训练判别器模块操作于由冻结的、经REPA预训练的SiT骨干从干净真实和生成潜在特征中提取的语义丰富特征。所有训练均在潜在空间中进行,无需图像解码。在精炼过程中,FiST消费其自身的中间预测,端点梯度穿过每个生成阶段。对于ImageNet上$256\ imes256$的类条件生成,我们的方法在三个阶段下实现了FID 1.11(IS 282),在两个阶段下实现了FID 1.15(IS 280)。这些结果表明,通过学习到的分布级监督,无需显式最小化Fréchet距离,即可实现具有竞争力的少步生成。消融实验表征了蒸馏、预训练检查点选择、精炼监督和跨阶段隐藏通信对生成质量的影响。
英文摘要
Iterative generation poses a joint optimization problem across steps, as intermediate predictions shape subsequent computations and ultimately determine the final output distribution. Few-step generators distilled from pretrained diffusion and flow-matching models make such optimization computationally practical end to end. We build on this opportunity with a distill-then-refine approach that uses teacher imitation to establish a strong initialization for a few-step rollout in latent space, then shifts to end-to-end refinement of the complete latent rollout against real data. We introduce FiST (Flow-in-Stage Transformer), an architecture that composes learned latent-state transitions in a few stages using a shared Transformer, with optional cross-stage hidden communication. Distillation applies teacher-forced regression to selected states along teacher trajectories; refinement replaces this supervision with adversarial and auxiliary classification objectives on the final latent output. A trainable discriminator module operates on semantically rich features extracted from clean real and generated latents by a frozen SiT backbone pretrained with REPA. All training takes place in latent space, without image decoding. During refinement, FiST consumes its own intermediate predictions, and endpoint gradients pass through every generation stage. For class-conditional generation on ImageNet at $256\times256$, our approach achieves FID 1.11 (IS 282) with three stages and FID 1.15 (IS 280) with two. These results demonstrate competitive few-step generation through learned distribution-level supervision, without explicit Fréchet-distance minimization. Ablations characterize how distillation, pretrained checkpoint choices, refinement supervision, and cross-stage hidden communication affect generation quality.