像素空间文本到图像扩散模型训练的实证研究
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
- Alibaba Token Hub(阿里巴巴令牌中心)
- Alibaba Group(阿里巴巴集团)
- Nanjing University(南京大学)
- University of California San Diego(加利福尼亚大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文通过实证研究提出潜空间到像素空间的训练策略,确定关键设计选择,使像素空间文本到图像扩散模型性能媲美或超越潜空间模型,且推理加速3.18至4.75倍,为相关研究提供实用指南。
AI中文摘要:
本文研究生成建模中日益重要的像素空间扩散模型主题。尽管已有大量研究探索该主题,但多数聚焦于小规模或类条件设置,因此仍缺乏可媲美或超越成熟潜空间对应模型的像素空间模型训练实用方案。通过全面实证研究,我们首先发现像素空间的大规模直接预训练收敛速度显著慢于潜空间。该观察促使我们提出一种潜空间到像素空间的策略:先在潜空间高效获取生成先验,再在训练后阶段过渡至像素空间。随后,我们系统研究了决定该过渡的关键设计选择,包括权重初始化、数据组成、预测目标、解码器架构和噪声调度,确定了一套实用方案,使所得像素空间模型可媲美或超越潜空间对应模型,同时实现3.18至4.75倍的端到端推理加速。我们希望研究结果为未来像素空间生成研究提供有用的实证见解和实用指南。
英文摘要:
This paper investigates an increasingly important topic in generative modeling: pixel-space diffusion models. Although numerous studies have explored this topic, most focus on small-scale or class-conditional settings. Consequently, a practical recipe for training pixel-space models that rival or exceed well-established latent-space counterparts remains elusive. Through a comprehensive empirical study, we first observe that direct large-scale pre-training in pixel space converges substantially more slowly than in latent space. This observation motivates a latent-to-pixel strategy that acquires generative priors efficiently in latent space and transitions to pixel space during post-training. We then systematically investigate the key design choices governing this transition, including weight initialization, data composition, prediction target, decoder architecture, and noise schedule, and identify a practical recipe that makes the resulting pixel-space models match or outperform their latent-space counterparts while delivering 3.18 to 4.75 times end-to-end inference speedups. We hope that our findings provide useful empirical insights and practical guidelines for future research on pixel-space generation.