解锁少步扩散以实现忠实预览
Unlocking Few-Step Diffusion for Faithful Previews
浏览论文内容
中文总结 AI 辅助
本文发现通过优化初始噪声,冻结的少步采样器可复现全步输出,并提出基于端点监督的输入修正方法,显著提升预览保真度与候选筛选效率。
中文摘要 AI 辅助
在扩散工作流中,采样延迟会不断累积,因为用户通常会生成并丢弃许多候选样本,最终才保留一个。令人惊讶的是,标准少步采样器的较差输出并非反映重建能力的缺失:通过仅优化初始噪声,冻结的3-4步采样器能够紧密复现其对应的全步输出。基于这一发现,我们利用端点监督学习对初始噪声和去噪更新的修正,从而改善与由相同噪声和提示生成的全步输出的一致性。由此产生的预览使用户能够低成本地筛选候选样本,并将全步生成保留给有前景的候选。输入修正还能在无需重新训练的情况下跨采样预算迁移。实验表明,在参考保真度方面有显著提升,包括在无条件基准上比重新训练的LD3降低53-78%的重建均方误差,同时在SD1.5、SDXL和FLUX.1-dev上改善了排序保持和候选选择。
英文摘要
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
发表机构
- Rutgers University(罗格斯大学)
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。