发表机构
The Hong Kong Polytechnic University; OPPO Research Institute(香港理工大学; OPPO研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LDM-is-AE,将潜扩散模型视为自编码器,通过拆分DiT为编码和解码组件并施加图像空间监督,实现端到端单阶段训练,在256和512分辨率下FID分别达1.80和1.90。
AI 中文摘要
潜扩散模型(LDMs)通常采用两阶段流程:首先预训练一个自编码器(AE)来定义潜空间,然后训练一个扩散模型在其中进行去噪。这种两阶段设计引入了表示不匹配,因为潜空间是针对重建而非适应去噪动态而优化的。我们揭示LDM本身就是一个自编码器,并由此提出LDM-is-AE,一种端到端的单阶段LDM训练框架,无需单独训练的分词器。我们的关键观察是,LDM骨干网络在每一步去噪中实际上执行了从潜变量到特征再到潜变量的变换,这可以解释为内部的解码-编码过程。利用这一结构,我们将DiT骨干网络拆分为两个互补组件,DiT-E(即DiT编码)和DiT-D(即DiT解码),并在所有时间步对中间特征施加图像空间监督。我们的模型鼓励内部表示在整个去噪过程中与图像域对齐,从而建立显式的潜变量-图像-潜变量路径。在零噪声时间步,我们的模型进一步执行图像到潜变量再到图像的映射,对应于自编码过程。因此,LDM-is-AE以端到端方式联合学习潜表示和去噪动态,产生针对生成过程定制的扩散原生潜空间。实验表明,LDM-is-AE展现出极具竞争力的生成性能,在256x256和512x512类别条件图像生成上分别实现了1.80和1.90的FID。
英文摘要
Latent Diffusion Models (LDMs) typically adopt a two-stage pipeline: an auto-encoder (AE) is first pre-trained to define a latent space, then a diffusion model is trained to perform denoising within it. Such a two-stage design introduces a representation mismatch, as the latent space is optimized for reconstruction rather than adapting the denoising dynamics. We reveal that the LDM itself is an AE, and consequently present LDM-is-AE, an end-to-end one-stage LDM training framework that eliminates the need for a separately trained tokenizer. Our key observation is that the LDM backbone actually performs a latent-to-feature-to-latent transformation at each denoising step, which can be interpreted as an internal decoding--encoding process. Leveraging this structure, we split the DiT backbone into two reciprocal components, DiT-E (i.e., DiT Encoding) and DiT-D (i.e., DiT Decoding), and impose image-space supervision on the intermediate features across all timesteps. Our model encourages the internal representation to align with the image domain throughout denoising, thereby establishing an explicit latent-to-image-to-latent path. At the zero-noise timestep, our model further performs an image-to-latent-to-image mapping, corresponding to an auto-encoding process. As a result, LDM-is-AE jointly learns latent representations and denoising dynamics in an end-to-end manner, yielding a diffusion-native latent space tailored to the generation process. Experiments demonstrate that LDM-is-AE exhibits highly competitive generation performance, achieving an FID of 1.80 and 1.90 on 256x256 and 512x512 class-conditional image generation, respectively.
CommentsAccepted by NIPS 2026. More info can be found in https://github.com/PolyU-VCLab/LDMisAE