arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03641cs.LG

IDRF:掩码离散扩散模型的逆蒸馏奖励微调

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

发表机构应用人工智能研究所 · 穆罕默德·本·扎耶德人工智能大学 · 巴黎综合理工学院应用数学中心
另 2 家 · 查看机构详情
  • Applied AI Institute(应用人工智能研究所)
  • MBZUAI(穆罕默德·本·扎耶德人工智能大学)
  • CMAP, École Polytechnique(巴黎综合理工学院应用数学中心)
  • Institute of Foundation Models(基础模型研究所)
  • LRE, EPITA(EPITA 研究与工程实验室)

机构由 AI 辅助整理,请以论文原文为准。

Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin

首次发表
浏览论文内容

中文总结 AI 辅助

IDRF通过逆蒸馏正则化替代序列级KL惩罚,实现少步掩码离散扩散模型的奖励微调,在DNA、图像和文本生成中减少32倍去噪步骤,同时保持高质量。

中文摘要 AI 辅助

掩码离散扩散模型为自回归生成提供了一种有前景的替代方案,但迭代采样可能成本高昂,且难以处理的序列似然性使奖励微调复杂化。我们引入了IDRF,一个用于少步掩码离散扩散生成器奖励微调的框架。从标准的反向KL正则化目标出发,IDRF用逆蒸馏正则化替代了难以处理的序列级KL惩罚。借助最优辅助去噪器,我们证明了总体逆蒸馏损失上界于参考分布的序列级KL散度。IDRF优化了该损失的基于轨迹的替代形式,无需参考模型展开,因此学生模型保留其自身的少步采样器。我们将少步生成视为有限时域马尔可夫决策过程,并通过学生轨迹上的裁剪策略梯度目标优化奖励。在DNA、图像和文本生成中,IDRF实现了高奖励,去噪步骤比参考模型减少多达32倍,同时缓解了奖励黑客攻击并保持了样本质量。

英文摘要

Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.

↑