arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAST:用于高效扩散模型的提示自适应采样终止

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

Renye Yan, Jikang Cheng, You Wu, Wei Peng, Zongwei Wang, Ling Liang, Yimao Cai

arXiv 2608.06794首次发表:更新:

发表机构

Peking University; Nanjing University; Stanford University(北京大学; 南京大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PAST是一种提示自适应采样终止方法,通过双自适应协调机制提升扩散模型RL微调的计算效率与偏好优化质量,效率最高提升66.7%,质量最高提升29.5%。

AI 中文摘要

尽管扩散模型在文本到图像任务中已取得显著进展,但直接优化下游目标时仍存在局限性。强化学习(RL)虽能实现针对性优化,但现有方法普遍受限于微调效率低和奖励稀疏的问题。为应对这些挑战,我们提出PAST,该方法通过联合感知去噪进度与提示难度,提供差异化奖励并自适应调整训练回合长度。具体而言,我们设计了内在奖励范式以补偿稀疏的外在奖励,并引导模型探索能更高效脱离噪声模式的路径,同时为内在奖励提供理论依据。随后,PAST动态监测去噪完成度以及图像结构与提示语义之间的语义对齐度,当两个指标满足生成要求时,系统自适应终止训练,从而根据提示难度和当前生成过程合理分配回合长度。最后,基于预测的残差噪声水平,我们建立了双自适应协调机制,该机制不仅能平衡外在奖励与内在奖励,还能平衡探索与收敛。实验结果表明,PAST通过其双自适应调节机制,可将现有RL微调方法的计算效率提升最高66.7%,同时将偏好优化质量提升最高29.5%。

英文摘要

While diffusion models have made significant progress in text-to-image tasks, they still exhibit limitations when directly optimizing downstream objectives. Although Reinforcement Learning (RL) enables targeted optimization, existing methods are generally constrained by low-efficiency fine-tuning and sparse rewards. To address these challenges, we propose PAST, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty. Specifically, we design an intrinsic reward paradigm to compensate for sparse extrinsic rewards and guide the model to explore paths that diverge more efficiently from noise patterns. We further provide theoretical justification for intrinsic rewards. Then, PAST dynamically monitors denoising completion and semantic alignment between image structures and prompt semantics. When both metrics satisfy generation requirements, the system adaptively terminates training. This enables appropriate allocation of episode lengths based on prompt difficulty and the current generation process. Finally, based on the predicted residual noise level, we establish a dual adaptive coordination mechanism. Specifically, it not only balances the extrinsic and intrinsic rewards but also balances the exploration and convergence. Experimental results demonstrate that PAST enhances computational efficiency of existing RL fine-tuning methods by up to 66.7%, while improving preference optimization quality by up to 29.5% through its dual adaptive regulation mechanism.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑