arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PEAK:基于k稀疏自编码器(kSAEs)的精确且持久的概念擦除

PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders

Man Jiang, Ouxiang Li, Weibao Xue, Zhenhua Tang, Yuan Wang, Shuo Wang, Yanbin Hao

arXiv 2608.10985首次发表:更新:

AI 中文总结

本研究针对文本到图像扩散模型的概念擦除难题,提出基于k稀疏自编码器的PEAK框架,实现了精确持久的概念擦除,在I2P基准上大幅降低了冒犯性内容检测量与攻击成功率,同时保留了生成质量。

AI 中文摘要

由于对版权侵权、隐私侵犯和冒犯性内容的担忧日益加剧,从大规模文本到图像(T2I)扩散模型中擦除概念变得愈发关键。现有方法难以同时实现精确且持久的概念擦除:与概念相关的表示定位不准确可能会导致意外的语义干扰,而底层概念知识的不完全移除则会使概念可被对抗性恢复。为解决这一困境,我们提出了PEAK,一种基于k稀疏自编码器(kSAEs)的精确且持久的概念擦除框架。PEAK首先在扩散去噪网络的内部激活上训练kSAE,以将密集表示分解为可解释的稀疏特征。通过对比目标提示和非目标提示诱导的稀疏激活,PEAK根据激活强度和频率识别出一组紧凑的目标特定特征。这些定位后的特征随后被用于参数优化,其中PEAK选择性地抑制与目标相关的激活,同时保留与原始模型互补的非目标激活。这种特征引导的优化将概念擦除直接嵌入到扩散参数中,无需额外的推理时干预,并有助于实现针对对抗性攻击的有效持久性。大量实验表明,PEAK能实现有效且鲁棒的概念擦除:在I2P基准上,PEAK将NudeNet检测量从582降至6,将平均攻击成功率(ASR)从96.52%降至5.63%,并在MS-COCO上保留了接近零KID的通用生成质量。我们的代码和模型可在以下网址获取:this https URL

英文摘要

Erasing concepts from large-scale text-to-image (T2I) diffusion models has become increasingly crucial due to the growing concerns over copyright infringement, privacy violations, and offensive content. Existing approaches struggle to achieve both precise and persistent concept erasure: inaccurate localization of concept-related representations may cause unintended semantic interference, while incomplete removal of the underlying concept knowledge allows adversarial recovery. To address this dilemma, we propose PEAK, a \textbf{\textit{precise}} and \textbf{\textit{persistent}} concept erasure framework via k-Sparse Autoencoders (kSAEs). PEAK first trains a kSAE on internal activations of the diffusion denoising network to decompose dense representations into interpretable sparse features. By contrasting sparse activations induced by target and non-target prompts, PEAK identifies a compact set of target-specific features according to both activation strength and frequency. These localized features are then used for parameter optimization, where PEAK selectively suppresses target-related activations while preserving complementary non-target ones towards the original model. This feature-guided optimization embeds concept erasure directly into diffusion parameters, eliminating the need for additional inference-time intervention and facilitating effective persistence against adversarial attacks. Extensive experiments demonstrate that PEAK achieves effective and robust concept erasure. On the I2P benchmark, PEAK reduces NudeNet detections from 582 to 6, lowers the average attack success rate (ASR) from 96.52\% to 5.63\%, and preserves general generation quality on MS-COCO with a near-zero KID. Our code and models are available at: https://github.com/manmanTAT/PEAK

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑