arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15341cs.CV

TEA:用于文本到图像模型中稳健概念擦除的文本编码器对齐

TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models

  • Mila – Quebec AI Institute(米拉-魁北克人工智能研究所)
  • McGill University(麦吉尔大学)
  • Google Research(谷歌研究院)

机构由 AI 辅助整理,请以论文原文为准。

Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh, Golnoosh Farnadi

AI总结:

该研究提出轻量级文本编码器对齐框架 TEA,将概念擦除建模为文本表示空间的域对齐问题,仅微调文本编码器,在零推理开销下提升了文本到图像模型对对抗攻击的概念擦除稳健性,且模型无关,在 Stable Diffusion 系列模型上表现优异。

AI中文摘要:

文本到图像扩散模型可能被滥用,通过绕过内置安全机制的对抗性提示或 paraphrased(意译)提示生成有害内容。现有概念擦除方法往往存在对对抗性提示的稳健性有限、良性生成质量下降,或依赖引入持续计算开销的推理时干预等问题。为解决这些局限,我们将概念擦除建模为文本表示空间中的域对齐问题。我们提出轻量级文本编码器对齐框架(TEA),仅对文本编码器进行微调,同时保持生成主干完全冻结。给定概念-锚定提示对,我们的方法训练一个判别器,以区分含概念提示的词级表示与安全锚定提示的词级表示,同时更新文本编码器使这些表示不可区分。TEA 引入零推理时开销,且仅需少量微调步骤,因此大规模部署时效率极高。尽管效率高,TEA 在 Stable Diffusion v1.4 上对黑盒和白盒对抗攻击实现了最先进的擦除稳健性,同时保留了对良性提示的生成质量。此外,TEA 是模型无关的,在 Stable Diffusion v3.5 上实现了最低的攻击成功率,将概念擦除扩展到带有 T5 条件的 Rectified Flow Transformer 架构,而此前的方法在该架构上基本未被探索。代码可在 https://github.com/... (注:原链接为 this https URL)获取。

英文摘要:

Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasure methods often suffer from limited robustness against adversarial prompts, degradation of benign generation quality, or reliance on inference-time interventions that introduce persistent computational overhead. To address these limitations, we formulate concept erasure as a domain alignment problem in the text representation space. We propose a lightweight Text Encoder Alignment framework (TEA) that fine-tunes only the text encoder while keeping the generative backbone fully frozen. Given concept--anchor prompt pairs, our method trains a discriminator to distinguish token-level representations of concept-containing prompts from those of safe anchor prompts, while updating the text encoder to make these representations indistinguishable. TEA introduces zero inference-time overhead and requires only a small number of fine-tuning steps, making it highly efficient to deploy at scale. Despite this efficiency, TEA achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts. Furthermore, TEA is model-agnostic and achieves the lowest attack success rate on Stable Diffusion v3.5, extending concept erasure to a Rectified Flow Transformer architecture with T5 conditioning where prior methods remain largely unexplored. Code is available at \href{https://github.com/alirezafarashah/TEA.git}{https://github.com/alirezafarashah/TEA.git}

↑