arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34406cs.LGcs.CVstat.ML

解锁少步扩散以实现忠实预览

Unlocking Few-Step Diffusion for Faithful Previews

Jing Jia, Sifan Liu, Guanyang Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本文发现通过优化初始噪声,冻结的少步采样器可复现全步输出,并提出基于端点监督的输入修正方法,显著提升预览保真度与候选筛选效率。

中文摘要 AI 辅助

在扩散工作流中,采样延迟会不断累积,因为用户通常会生成并丢弃许多候选样本,最终才保留一个。令人惊讶的是,标准少步采样器的较差输出并非反映重建能力的缺失:通过仅优化初始噪声,冻结的3-4步采样器能够紧密复现其对应的全步输出。基于这一发现,我们利用端点监督学习对初始噪声和去噪更新的修正,从而改善与由相同噪声和提示生成的全步输出的一致性。由此产生的预览使用户能够低成本地筛选候选样本,并将全步生成保留给有前景的候选。输入修正还能在无需重新训练的情况下跨采样预算迁移。实验表明,在参考保真度方面有显著提升,包括在无条件基准上比重新训练的LD3降低53-78%的重建均方误差,同时在SD1.5、SDXL和FLUX.1-dev上改善了排序保持和候选选择。

英文摘要

Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.

发表机构

  • Rutgers University(罗格斯大学)
  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

↑