arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SALD:扩散模型的自引用优势学习

SALD: Self-Referenced Advantage Learning for Diffusion Models

Aryan Das, Surjo Dey, Koushik Biswas, Swalpa Kumar Roy, Moloud Abdar, Arnab Bhattacharya, Vinay Kumar Verma

arXiv 2610.01496首次发表:更新:

发表机构

Indian Institute of Technology Kanpur; Rajiv Gandhi Institute of Petroleum Technology; IIIT Delhi; Tezpur University; The University of Queensland(印度理工学院坎普尔分校; 拉吉夫·甘地石油技术学院; 德里印度信息技术学院; 特兹普尔大学; 昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SALD提出一种自引用训练框架,通过比较同一模型在不同噪声路径下的误差,利用优势引导扩散、时间记忆和频谱分解来提升扩散模型的生成质量,无需外部教师网络或额外参数。

AI 中文摘要

近期关于语言模型适配的研究表明,单个模型可以通过在演示或反馈增强的上下文中评估自身行为来获得信息丰富的训练信号,这得益于由学生模型学习参数驱动的教师网络。受这一内部参考原理的启发,我们研究了扩散模型如何在无需外部演示或教师网络的情况下识别自引用训练信号。我们引入了SALD,一个自引用训练框架,它使用同一模型在两个噪声级别下评估每个图像-文本对。较容易的低噪声路径在不进行梯度跟踪的情况下进行评估以提供参考,而较困难的高噪声路径则提供训练梯度。SALD并非直接蒸馏易路径的预测,而是利用两条路径误差之间的差异来调整难路径的目标。所提出的优势引导扩散(AGD)将此相对误差转化为可微分的样本级权重。时间优势记忆(TAM)在训练过程中累积相对难度,并调整未来两个噪声级别之间的差距。频谱优势分解(SAD)进一步比较两条路径的残差功率谱,并构建一个可微分的、基于频率的潜在元素权重。所有组件共享一组模型参数,在训练或推理期间既不需要外部教师网络,也不需要额外的可训练参数,且无需修改推理过程。在多种架构和数据集上的实验表明,生成质量持续提升,而组件级消融实验量化了所提出组件的贡献。

英文摘要

Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student's learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑