高效文本到图像生成:扩散模型的自适应步数调度控制器
Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models
浏览论文内容
中文总结 AI 辅助
针对文本到图像扩散模型固定步数效率低的问题,提出自适应步数调度控制器,无需训练,通过混合调度与误差评估动态调整步数,在COCO和DiffusionDB上减少推理时间并保持视觉保真度。
中文摘要 AI 辅助
文本到图像的扩散模型通常使用固定数量的去噪步骤,以在时间成本和图像质量之间取得平衡。然而,最优步骤数量取决于输入文本提示的复杂性。我们提出了一种自适应扩散控制器,无需额外模型训练即可动态调整步骤数量,以高效生成高质量图像。通过利用具有不同步长的混合步数调度,并在每个时间步评估误差项差异,我们的方法在调度之间进行转换以优化性能。在COCO和DiffusionDB上的实验表明,我们的方法在保持视觉保真度的同时减少了推理时间,为文本到图像扩散模型提供了一种更高效的替代方案。
英文摘要
Text-to-image diffusion models often use a fixed number of denoising steps, balancing time costs and image quality. However, the optimal number of steps depends on the complexity of the input text prompt. We propose an adaptive diffusion controller that dynamically adjusts the number of steps to generate high-quality images efficiently, without additional model training. By leveraging a mixture of step schedules with varying step sizes and evaluating the error term discrepancy at each timestep, our method transitions between schedules to optimize performance. Experiments on COCO and DiffusionDB show that our approach reduces inference time while maintaining visual fidelity, offering a more efficient alternative for text-to-image diffusion models.
发表机构
- SAP(SAP公司)
- National University of Singapore(新加坡国立大学)
- Institute for Infocomm Research (I2R), A*STAR(资讯通信研究院(I2R),新加坡科技研究局)
机构由 AI 辅助整理,请以论文原文为准。