arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31882cs.LGcs.AI

DOHF:基于Doob的$h$-变换引导的在线扩散微调

DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance

Zhengyi Guo, Jiayuan Sheng, Wenpin Tang, David D. Yao

首次发表
浏览论文内容

中文总结 AI 辅助

提出DOHF方法,将Doob的h-变换转化为在线训练算法,通过最优性权重和局部校正估计蒸馏实现奖励微调,统一了多种引导方法,适用于黑盒奖励,提升生成对齐效果。

中文摘要 AI 辅助

基于奖励的扩散微调在面对理想结果稀少或条件校正成本高昂时面临实际挑战。在这项工作中,我们提出了扩散在线$h$-引导微调(DOHF),它将Doob的$h$-变换转化为一种实用的在线训练算法。DOHF为生成的样本分配最优性权重,在当前策略下估计归一化的局部校正$\nabla\log h$,并将其直接蒸馏到生成模型中。理论上,我们通过统一的$h$-变换视角刻画了群体最优的DiffusionNFT更新以及各种分类器自由引导方法。在方法上,我们的框架无需额外的网络评估即可适应黑盒和非可微奖励。我们进一步在三个经验场景中展示了改进的对齐效果。我们的工作展示了如何通过廉价估计和迭代蒸馏来适应概率条件,从而改进统计采样和视觉生成中的生成学习。

英文摘要

Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.

↑