发表机构
University of Science and Technology of China; HiDream.ai Inc.(中国科学技术大学; HiDream.ai公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对文本到视频扩散模型的安全版权问题,提出基于稀疏自编码器的EraseSAE框架,通过分解-归因-擦除流程实现细粒度概念擦除,效果优于现有方法且质量损失极小。
AI 中文摘要
近期文本到视频(T2V)扩散模型的进展展现出卓越的生成能力,但它们依赖整理松散的训练数据引发了紧迫的安全与版权问题。概念擦除提供了一种原则性解决方案,可从预训练模型中移除不需要的语义同时保留其余概念。然而现有方法通常以粗粒度操作,与概念表示的细粒度、分布式特性不匹配,导致擦除不完整或生成质量下降。我们认为手术级擦除根本上需要在单语义特征层面进行干预,其中每个单元编码一个可解释的单一概念。为此,我们提出EraseSAE,这是一种新框架,利用稀疏自编码器通过原则性的分解-归因-擦除流程在基于DiT的T2V扩散模型中实现手术级概念擦除。我们首先引入分区卷积稀疏自编码器,它将密集的时空激活分解为解耦、可解释的稀疏特征,同时保留时空一致性。对比归因机制随后对比配对提示的激活,以分离概念特定的特征核。推理时,从识别出的核导出的时间步解析时空掩码将擦除限制在目标概念活跃的区域,保留无关内容完整。在不同扩散模型和概念擦除任务上的大量实验表明,EraseSAE实现了精确且鲁棒的概念擦除,同时质量下降极小,显著优于最先进的方法。代码可在https URL获取。
英文摘要
Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative capabilities, yet their reliance on loosely curated training data raises pressing safety and copyright concerns. Concept erasure offers a principled remedy by removing unwanted semantics from pretrained models while preserving remaining concepts. However, existing approaches typically operate at a coarse granularity misaligned with the fine-grained, distributed nature of concept representations, leading to incomplete removal or degraded generation quality. We argue that surgical erasure fundamentally requires intervention at the level of monosemantic features, where each unit encodes a single interpretable concept. To this end, we propose EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attribute-erase pipeline. We first introduce the Partitioned Convolutional Sparse Autoencoder, which decomposes dense spatiotemporal activations into disentangled, interpretable sparse features while preserving spatiotemporal coherence. A contrastive attribution mechanism then contrasts activations from paired prompts to isolate concept-specific feature kernels. At inference, timestep-resolved spatiotemporal masks derived from the identified kernels confine erasure to regions where the target concept is active, leaving unrelated content intact. Extensive experiments across diverse diffusion models and concept erasure tasks demonstrate that EraseSAE achieves precise and robust concept removal with minimal quality degradation, substantially outperforming state-of-the-art methods. The code is available at https://github.com/HiDream-ai/EraseSAE.
CommentsAccepted to ECCV 2026