arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlowErase-OPD:基于流匹配模型中锚定在线策略蒸馏的多概念擦除方法

FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models

Yi Sun, Yimin Zhou, Xinhao Zhong, Zhiqi Zhang, Junhao Li, Bin Chen

arXiv 2608.07620首次发表:更新:

AI 中文总结

FlowErase-OPD是基于在线策略蒸馏的多概念擦除框架,通过引入AMTD和ARC缓解擦除与生成能力的权衡,在多概念擦除任务中实现了优于现有方法的性能,且模型鲁棒性较强。

AI 中文摘要

流匹配模型的最新进展大幅提升了文本到图像生成的质量,但也因可能生成有害或不合规内容引发了日益增长的安全担忧。现有针对流匹配模型的概念擦除方法主要聚焦于移除单个概念,而同时有效擦除多个概念仍具挑战性。我们提出FlowErase-OPD,这是一种基于在线策略蒸馏(OPD)的多概念擦除框架。我们的方法首先将多个单概念擦除模型蒸馏为一个统一的LoRA模块,并引入锚定多教师蒸馏(AMTD),该方法整合了一个保留教师以缓解概念擦除与生成能力保留之间的权衡。为进一步提升多个擦除目标的协调性,我们开发了自适应保留控制(ARC),其在整个训练过程中动态调整每个擦除教师的采样频率、损失权重,以及擦除教师与保留教师的相对贡献。针对裸体、物体和艺术风格擦除的大量实验表明,FlowErase-OPD在擦除效果、图像质量与语义对齐之间的权衡上持续得到改善,在多种多概念擦除场景中实现了最先进的性能。此外,得到的模型对对抗攻击表现出较强的鲁棒性。这些结果凸显了在线策略蒸馏作为流匹配模型中安全可控生成的原则性框架的潜力。

英文摘要

Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods for flow matching models predominantly focus on removing individual concepts, while effectively erasing multiple concepts simultaneously remains challenging. We propose FlowErase-OPD, a framework for multi-concept erasure based on on-policy distillation (OPD). Our approach first distills multiple single-concept erased models into a unified LoRA module and introduces Anchored Multi-Teacher Distillation (AMTD), which incorporates a retention teacher to mitigate the trade-off between concept erasure and preservation of generative capabilities. To further improve the coordination of multiple erasure objectives, we develop Adaptive Retention Control (ARC), which dynamically adjusts the sampling frequency and loss weight of each erasure teacher, together with the relative contribution of erasure and retention teachers throughout training. Extensive experiments on nudity, object, and artistic-style erasure demonstrate that FlowErase-OPD consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving state-of-the-art performance across diverse multi-concept erasure settings. Furthermore, the resulting models exhibit strong robustness against adversarial attacks. These results highlight the potential of on-policy distillation as a principled framework for safe and controllable generation in flow matching models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑