GenFirst:用于稳定端到端隐空间生成建模的生成优先于重构方法
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出GenFirst策略,实现无隐空间塌陷的端到端隐空间生成建模,在图像生成任务上取得优异指标,并将框架扩展至多模态生成场景。
AI中文摘要:
隐空间生成模型通常遵循两阶段流程:先训练用于重构的变分自编码器,再在冻结的隐空间上训练生成模型。由于针对重构优化的隐空间未必利于生成,联合训练两种模型是极具吸引力的替代方案,但直接端到端训练仍具挑战性,因为它易出现隐空间塌陷,且面临生成与重构的冲突。本文通过分析不同目标如何塑造隐空间重新审视该问题,得出两项关键见解:第一,Kullback-Leibler散度目标中的熵项对防止塌陷至关重要——重构与先验拟合会使后验收缩,而熵能保留非退化的隐空间不确定性;第二,重构与生成呈现不对称学习动态:重构速度快且受强监督,而生成速度慢且更难优化。基于这些见解,我们实现了首个无隐空间塌陷的直接端到端训练,并提出GenFirst,一种简单的生成优先于重构策略:生成目标先在弱重构压力下塑造隐空间,之后逐步强化重构以恢复视觉细节。我们采用带精确似然的连续自回归先验和带隐式似然的SiT先验对GenFirst进行验证,在ImageNet-256上,SiT在带CFG时的gFID为0.97,不带CFG时为1.45;MMDiT在文本到图像生成任务上的GenEval分数达0.90。除图像生成外,我们将该框架扩展至用于生成与表示学习的共享视觉隐空间,以及连续统一的文本-图像生成。这些结果证明了稳定端到端隐空间学习在不同生成先验和模态间的通用性。
英文摘要:
Latent generative models typically follow a two-stage pipeline, training a variational autoencoder for reconstruction and then a generative model on the frozen latent space. Since reconstruction-optimized latents are not necessarily generation-friendly, jointly training both models is an appealing alternative. However, direct end-to-end training remains challenging, as it is prone to latent collapse and faces a generation-reconstruction conflict. We revisit this problem by analyzing how different objectives shape the latent space and identify two key insights. First, the entropy term in the Kullback-Leibler divergence objective is essential for preventing collapse: reconstruction and prior fitting tend to shrink the posterior, while entropy preserves non-degenerate latent uncertainty. Second, reconstruction and generation exhibit asymmetric learning dynamics: reconstruction is fast and strongly supervised, whereas generation is slower and harder to optimize. Based on these insights, we achieve the first direct end-to-end training without latent collapse and propose GenFirst, a simple generation-before-reconstruction strategy. The generative objective first shapes the latent space under weak reconstruction pressure, after which reconstruction is progressively strengthened to recover visual details. We validate GenFirst with continuous autoregressive priors with exact likelihoods and SiT priors with implicit likelihoods. With our end-to-end objective and GenFirst, SiT achieves a gFID of 0.97 with CFG and 1.45 without CFG on ImageNet-256, while MMDiT reaches a GenEval score of 0.90 on text-to-image generation. Beyond image generation, we extend the framework to shared visual latents for generation and representation learning, and to continuous unified text-image generation. These results demonstrate the generality of stable end-to-end latent learning across generative priors and modalities.