SwiftExplorer:基于快速多样性探索的无训练扩散模型对齐
SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration
另 1 家 · 查看机构详情
- Peking University(北京大学)
- Nanjing University(南京大学)
- Huazhong Agricultural University(华中农业大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
SwiftExplorer通过继承-重启探索和质量-效率仲裁机制,在无训练条件下平衡扩散模型的多样性与对齐,降低计算成本并提升偏好、保真度等指标。
中文摘要 AI 辅助
扩散模型具有通用的生成能力,但难以与特定目标对齐。微调可以改善对齐,但其训练成本往往令人望而却步。这催生了无训练方法,这些方法在采样中应用目标引导项,将生成分布偏向指定区域,例如高奖励区域。然而,这些方法面临两个问题:(1)强烈的方向性偏差收窄了预训练分布和生成多样性;(2)不加区分的恒定引导无法剪除冗余信号,损害了质量和效率。为解决上述挑战,我们提出了SwiftExplorer,一个插件,它缓解了由过度多样性损失引起的分布坍缩,并降低了计算成本。首先,我们采用继承-重启探索机制以避免早期收敛,同时探索也增加了高奖励轨迹的可能性。此外,它平衡了多样性和保真度,在增加多样性的同时不会导致分布过度偏移。其次,我们的质量-效率仲裁机制通过移除错误信号来改进引导,并通过在完整性和边际奖励增益最优时动态停止生成来减少计算。在大量实验和不同类型的评估指标中,所提出的SwiftExplorer在所有指标上均取得了优异性能,包括偏好、保真度、多样性和丰富性。
英文摘要
Diffusion models have general generative abilities but struggle to align with specific objectives. Fine-tuning can improve alignment, yet its training cost is often prohibitive. This led to training-free methods that apply objective-guided terms in sampling to bias the generation distribution toward designated regions, e.g., high-reward areas. However, these methods face two issues: (1) the strong directional bias narrows the pretrained distribution and generation diversity, and (2) indiscriminate constant guidance fails to prune redundant signals, hurting both quality and efficiency. To address the above challenges, we propose SwiftExplorer, a plugin that mitigates distribution collapse caused by excessive diversity loss and reduces compute costs. First, we adopt an Inheritance-Restart exploration mechanism to avoid early convergence, while exploration also increases the likelihood of high-reward trajectories. Additionally, it balances diversity and fidelity, adding diversity without causing a distribution over-shift. Second, our Quality-Efficiency arbitration mechanism improves guidance by removing incorrect signals, and it reduces computation by dynamically stopping generation when completeness and marginal reward gain are optimal. In an extensive number of experiments and different types of evaluation metrics, the proposed SwiftExplorer achieves excellent performance on all metrics, including preference, fidelity, diversity, and richness.