arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10859cs.CVcs.LG

在少步蒸馏文本到图像扩散模型中实现偏好驱动的遗忘

Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Gaurav Patel, Jun Fang, Greg Ver Steeg, Qiang Qiu, Sravan Sripada

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对少步蒸馏文本到图像扩散模型的遗忘难题,提出偏好驱动的遗忘框架,改进直接偏好优化以适配少步生成特性,实现高效且有效的概念移除。

中文摘要 AI 辅助

文本到图像扩散模型正越来越多地被蒸馏为少步变体并部署以实现快速推理,然而它们生成有害或不合规内容的能力带来了重大安全风险。数据驱动的遗忘方法通过使用专门的遗忘目标微调模型权重来抑制目标生成,关键在于这些目标隐含依赖多步去噪动态,这一假设在少步蒸馏(FSD)模型中不成立,导致遗忘效果不佳。此外,在未蒸馏的基础模型上执行遗忘,随后重新蒸馏以获得遗忘的FSD模型会产生大量计算和时间开销,在许多场景中不实用。因此,我们提出一种偏好驱动的遗忘框架,将直接偏好优化(DPO)重新应用于扩散模型。我们表明,围绕噪声预测误差构建的标准DPO及其遗忘衍生方法,由于FSD模型的生成动态发生改变,难以迁移到FSD模型。为克服这一点,我们引入一种修改后的偏好优化公式,明确与少步生成特性对齐,能够在保留少步效率和对理想(非目标)能力的强保留的同时,直接在FSD模型中移除概念。我们主要在身份和NSFW(裸体)移除任务上评估我们的框架,并将我们的方法扩展到对象级遗忘。大量实验表明,该方法实现了一致且有效的遗忘,以及强的保留性能,确立了其作为FSD模型遗忘的实用且有原则的解决方案。

英文摘要

Text-to-image diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, enabling direct concept removal in FSD models while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.

发表机构

  • Amazon AGI(亚马逊AGI)
  • Purdue University(普渡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑