arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FADE:基于帧感知扩散Transformer的多概念擦除用于视频遗忘

FADE: Frame-Aware Diffusion-Transformer-based Multi-Concept Erasure for Video Unlearning

Yuchen Li, Kaiyuan Deng, Chaoran Feng, Zhenyu Tang, Li Yuan

arXiv 2610.03980首次发表:更新:

发表机构

Peking University; University of Arizona(北京大学; 亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出帧感知扩散擦除(FADE)框架,通过联合闭式编辑和帧感知低秩适配器解决视频概念擦除中的帧重新激活和多概念问题,在16概念擦除中显著降低残余准确率并保持整体性能。

AI 中文摘要

文本到视频(T2V)扩散模型可能再现受版权保护、暴力或露骨内容,这促使了概念擦除的研究:从预训练模型中移除指定概念,同时保持其在其他所有内容上的行为。现有的T2V擦除方法遗留了两个问题。其帧无关的抑制可能留下孤立帧,其中被擦除的概念会重新出现,这是一个片段级平均值所掩盖的帧重新激活缺口;并且它们通常一次只针对一个目标概念或类别进行评估。我们提出了帧感知扩散擦除(FADE),一个多概念视频遗忘框架。FADE首先应用联合闭式键/值编辑来抑制所有目标概念,然后训练每个概念的帧感知低秩适配器,其强度由帧索引和去噪时间步门控,以消除残余的逐帧泄漏。每个适配器使用其他目标的提示作为硬负样本进行训练,这保持了不同适配器的概念特定组件良好分离,并且一个基于相似性的软路由器根据提示组合适配器。在从单个Wan2.1-T2V-1.3B骨干中擦除16个概念(物体、艺术风格和裸体)时,FADE将物体基准上的残余准确率降低到4.9%,而八个基线中最强的为15.5%,同时保持VBench平均值与未编辑模型相差在0.9%以内。在VLM评判者和盲人人类研究下排名不变,并且相对于最强基线的优势延续到组合多个擦除概念的提示、同时擦除30个名人身份,以及Wan2.1-T2V-14B、CogVideoX-2B和HunyuanVideo-1.5。

英文摘要

Text-to-video (T2V) diffusion models can reproduce copyrighted, violent, or explicit content, which motivates concept erasure: removing designated concepts from a pretrained model while preserving its behavior on everything else. Existing T2V erasure methods leave two problems open. Their frame-agnostic suppression can leave isolated frames in which an erased concept resurfaces, a frame-reactivation gap that clip-level averages obscure; and they are usually evaluated with one target concept or category at a time. We propose Frame-Aware Diffusion Erasure (FADE), a multi-concept video unlearning framework. FADE first applies a joint closed-form key/value edit that suppresses all target concepts, then trains per-concept frame-aware low-rank adapters whose strength is gated by the frame index and the denoising timestep to remove residual per-frame leakage. Each adapter is trained with the other targets' prompts as hard negatives, which keeps the concept-specific components of different adapters well separated, and a similarity-based soft router combines the adapters according to the prompt. With 16 concepts (objects, artistic styles, and nudity) erased from a single Wan2.1-T2V-1.3B backbone, FADE reduces the residual accuracy on the object benchmark to 4.9%, against 15.5% for the strongest of eight baselines, while keeping the VBench average within 0.9% of the unedited model. The ranking is unchanged under a VLM judge and a blinded human study, and the advantage over the strongest baseline carries over to prompts that combine several erased concepts, to 30 simultaneously erased celebrity identities, and to Wan2.1-T2V-14B, CogVideoX-2B, and HunyuanVideo-1.5.

Comments29 pages. Code: https://github.com/INTOTHEMILD/FADE

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑