面向不确定性引导的扩散模型后训练的样本自适应潜在奖励
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
浏览论文内容
中文总结 AI 辅助
本文提出\textsc{SURE}框架,含\textsc{SURE-LRM}与\textsc{SURE-REFL},通过样本自适应潜在奖励与不确定性引导反馈优化扩散模型,在偏好预测、SOTA性能及VBench指标上表现优异。
中文摘要 AI 辅助
潜在奖励模型可在不将中间状态解码到像素空间的前提下监督视觉扩散模型,这使与人类偏好的对齐更高效,但现有潜在奖励模型仅输出标量分数,未估计每个预测的不确定性,导致生成器无法判断反馈是否可靠,可能使优化方向错误并引发奖励黑客行为。本文提出统一的图像与视频扩散模型潜在空间框架\textsc{SURE},用于学习奖励分布并直接利用其可靠性指导密集后训练。首先提出样本自适应潜在奖励模型\textsc{SURE-LRM},为每个带噪潜在预测高斯效用,均值预测奖励分数,方差反映无人工标注时的预测不确定性;学习到的分布随后通过不确定性引导的奖励反馈学习\textsc{SURE-REFL}指导后训练,该方法沿去噪轨迹提供不确定性引导的密集反馈,在选定的转换点处,\textsc{SURE-REFL}查询冻结的\textsc{SURE-LRM},将分离的方差转换为相同转换点处样本的可靠性权重,每个加权奖励仅通过其局部转换反向传播,整个过程保持在潜在空间,无需像素空间解码或完整去噪图。实验表明,\textsc{SURE-LRM}在偏好预测上优于强基线,\textsc{SURE-REFL}在多种指标上达到SOTA性能,进一步提升优化稳定性,且在评估方法中取得最高的VBench质量、语义及总得分。
英文摘要
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar scores. They do not estimate the uncertainty of each prediction. The generator therefore cannot determine which feedback is reliable. This can drive optimization in the wrong direction and lead to reward hacking. We propose \textsc{SURE}, a unified latent-space framework for image and video diffusion models. It learns reward distributions and directly uses their reliability to guide dense post-training. First, we propose sample-adaptive latent reward model (\textsc{SURE-LRM}). It predicts a Gaussian utility for each noisy latent. Its mean predicts the reward score. Its variance reflect the uncertainty of prediction without human annotation. The learned distribution then guides post-training through uncertainty-guided reward feedback learning (\textsc{SURE-REFL}). This method provides uncertainty-guided dense feedback along the denoising trajectory. At selected transitions, \textsc{SURE-REFL} queries the frozen \textsc{SURE-LRM}. It converts detached variance into reliability weights for samples at the same transition. Each weighted reward is backpropagated only through its local transition. The entire process remains in latent space and requires neither pixel-space decoding nor the full denoising graph. Experiments show that \textsc{SURE-LRM} improves preference prediction over strong baselines. \textsc{SURE-REFL} achieves the sota performance among various metrics and further improves optimization stability. It also achieves the highest VBench quality, semantic, and total scores among the evaluated methods.