arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习扩散模型的采样参数

Learning Sampling Parameters for Diffusion Models

Arisrei Lim, Yossi Gandelsman

arXiv 2607.23488首次发表:更新:

AI 中文总结

研究针对扩散模型推理时采样参数固定的问题,提出LeSAMP框架,将参数选择转化为强化学习问题,通过人类偏好模型和VLM评判优化,实验表明该方法相比基线有更高胜率,为改进扩散模型输出提供补充途径。

AI 中文摘要

文本到图像的扩散模型有许多推理时的采样参数,如提示、负提示、无分类器指导尺度和噪声时间表。这些参数通常手动选择一次并在不同提示和去噪时间步保持固定,尽管不同提示和生成阶段可能受益于不同参数值。我们引入LeSAMP框架来学习提示条件下、随时间步变化的采样参数,将参数选择表述为强化学习问题,用人类偏好模型和VLM作为评判的奖励来优化模型。在Flux.1[dev]和Stable Diffusion 3.5上评估发现,与基线相比,LeSAMP在使用人类偏好分数时胜率高达68.12%,使用VLM作为评判时胜率为73.37%,在用户研究中比之前基线胜率高达59.46%。结果表明学习到的采样参数策略为改进扩散模型输出提供了补充方法。

英文摘要

Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps, even though different prompts and stages of generation can benefit from different parameter values. We introduce LeSAMP, a framework for learning prompt-conditioned, timestep-varying sampling parameters. We formulate parameter selection as a reinforcement learning problem: Given a user prompt, a large language model is trained to emit schedules for the chosen sampling parameters. We optimize our model using rewards from human preference models and VLM-as-a-judge. We evaluate our model on Flux.1 [dev] and Stable Diffusion 3.5, and find that compared to baselines, LeSAMP has a win rate of up to 68.12% using human preference scores and 73.37% using VLM-as-a-judge. These gains are validated in a user study where we achieve win rates of up to 59.46% over previous baselines. Our results suggest that learned sampling-parameter policies provide a complementary approach to existing post-training methods for improving diffusion model outputs.

CommentsCode will be released

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑