AI 中文总结
研究如何在推理时优化初始噪声以恢复几步蒸馏扩散模型的提示多样性。提出MoNO方法,在低维噪声流形上进行流形约束噪声优化,能大步长更新且无需辅助目标,实验表明该方法可保持图像质量并提升提示多样性。
AI 中文摘要
几步蒸馏扩散模型能快速生成高质量图像,但往往会失去每个提示的多样性,在不同随机种子下生成近乎相同的样本。在推理时优化初始噪声是恢复这种多样性的一种有吸引力的方法,但现有方法在无约束的欧几里得空间中直接更新初始噪声,忽略了高斯先验的几何结构和模型对噪声频率的敏感性。因此,它们引入辅助质量控制目标来维持生成保真度,增加了计算量和加权超参数,同时仍需要保守更新以防止质量下降。在这项工作中,我们提出了MoNO,一种无需训练的方法,在低维、质量稳定的噪声流形上执行流形约束噪声优化。MoNO依次优化每个新的初始噪声,使其预测的视觉特征补充先前的生成结果,而仿射低频球面上的黎曼更新保留了先验似然性,并通过构造修复了不稳定的高频分量。这使得可以进行大步长的测地线更新,无需辅助质量控制目标,并且比先前的噪声优化方法在更少的迭代中收敛。对多个蒸馏文本到图像扩散模型的实验表明,MoNO在保持图像质量的同时,持续提高了每个提示的多样性。
英文摘要
Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across random seeds. Optimizing the initial noise at inference time offers an appealing way to recover this diversity, yet existing methods directly update the initial noise in an unconstrained Euclidean space, ignoring both the geometry of the Gaussian prior and the model's sensitivity to noise frequencies. They therefore introduce auxiliary quality-control objectives to maintain generation fidelity, adding compute and weighting hyperparameters while still requiring conservative updates to prevent degradation. In this work, we propose MoNO, a training-free method that performs Manifold-constrained Noise Optimization on a low-dimensional, quality-stabilizing noise manifold. MoNO sequentially optimizes each new initial noise so that its predicted visual feature complements previous generations, while Riemannian updates on an affine low-frequency sphere preserve prior likelihood and fix unstable high-frequency components by construction. This enables large geodesic steps, removes the need for auxiliary quality-control objectives, and converges in far fewer iterations than prior noise-optimization methods. Experiments with multiple distilled text-to-image diffusion models show that MoNO consistently improves per-prompt diversity while maintaining image quality.
Comments19 pages, 14 figures, 4 tables