IDRF:掩码离散扩散模型的逆蒸馏奖励微调
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
另 2 家 · 查看机构详情
- Applied AI Institute(应用人工智能研究所)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
- CMAP, École Polytechnique(巴黎综合理工学院应用数学中心)
- Institute of Foundation Models(基础模型研究所)
- LRE, EPITA(EPITA 研究与工程实验室)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
IDRF通过逆蒸馏正则化替代序列级KL惩罚,实现少步掩码离散扩散模型的奖励微调,在DNA、图像和文本生成中减少32倍去噪步骤,同时保持高质量。
中文摘要 AI 辅助
掩码离散扩散模型为自回归生成提供了一种有前景的替代方案,但迭代采样可能成本高昂,且难以处理的序列似然性使奖励微调复杂化。我们引入了IDRF,一个用于少步掩码离散扩散生成器奖励微调的框架。从标准的反向KL正则化目标出发,IDRF用逆蒸馏正则化替代了难以处理的序列级KL惩罚。借助最优辅助去噪器,我们证明了总体逆蒸馏损失上界于参考分布的序列级KL散度。IDRF优化了该损失的基于轨迹的替代形式,无需参考模型展开,因此学生模型保留其自身的少步采样器。我们将少步生成视为有限时域马尔可夫决策过程,并通过学生轨迹上的裁剪策略梯度目标优化奖励。在DNA、图像和文本生成中,IDRF实现了高奖励,去噪步骤比参考模型减少多达32倍,同时缓解了奖励黑客攻击并保持了样本质量。
英文摘要
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.