发表机构
ETH Zurich; Google(苏黎世联邦理工学院; 谷歌)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对扩散模型难以达到罕见模式的问题,提出方差校正时间偏移方法,并结合温度采样,无需重新训练就能提升样本多样性,在多个模型中以低成本实现质量和保真度提升,还能实现从粗到细的控制。
AI 中文摘要
扩散模型能忠实地再现训练分布,但也继承了其不平衡性,使得难以达到罕见或代表性不足的模式。一种自然的推理时补救方法是从高温目标\(p^{(\gamma)}_0(x) \propto p_0(x)^{\gamma}\)(\(0 < \gamma < 1\))进行采样,这会使主导模式变平坦并提升罕见模式。然而,简单的分数缩放虽能正确重新加权模式,但也会增加每个模式的方差,破坏反向扩散过程并降低样本质量。我们引入方差校正时间偏移,这是一种无需训练的修正方法,在偏移的时间步查询网络并将所得分数乘以\(\gamma\),在保留模式重新加权的同时消除方差膨胀。这种校正将简单的温度采样转变为预训练扩散和流匹配主干的实用多样性旋钮,无需重新训练。我们证明,在DiT、Stable Diffusion和Motion Diffusion模型中,以最小的样本质量和条件保真度成本实现了一致的提升。我们还表明,温度干预的时机实现了从粗到细的控制:高噪声阶段驱动模式间的组合多样性,而低噪声阶段在固定组合下驱动局部外观变化。
英文摘要
Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach. A natural inference-time remedy is to sample from the high-temperature target $p^{(γ)}_0(x) \propto p_0(x)^γ$ for $0 < γ< 1$, which flattens dominant modes and lifts rare ones. However, naive score scaling while correctly reweighting modes also inflates the per-mode variance, breaking the reverse diffusion process and degrading sample quality. We introduce variance-corrective time shifting, a training-free fix that queries the network at a shifted timestep and scales the resulting score by $γ$, canceling the variance inflation while preserving the mode reweighting. The correction turns simple temperature sampling into a practical diversity knob for pretrained diffusion and flow-matching backbones with no retraining, and we demonstrate consistent gains at minimal cost to sample quality and condition fidelity across DiT, Stable Diffusion and Motion Diffusion models. We further show that the timing of the temperature intervention enables coarse-to-fine control: high-noise stages drive compositional diversity across modes, while low-noise stages drive local appearance variation under a fixed composition.
CommentsWebpage: https://peizhuoli.github.io/diversify-diffusion