arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视频扩散模型中的运动概念遗忘

Motion Concept Unlearning in Video Diffusion Models

Ping Liu, Chi Zhang

arXiv 2609.36832首次发表:更新:

发表机构

University of Nevada, Reno; National University of Singapore(内华达大学雷诺分校; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对T2V扩散模型中的运动概念擦除问题,提出无需训练的MUTE方法,通过令牌中和提取概念方向、空间门控和速度修正实现概念特异性、空间选择性与时间自然性,在多个模型上优于现有基线。

AI 中文摘要

文本到视频(T2V)扩散模型能够生成踢、刺、射击等动作的逼真描绘,这引发了安全问题,促使人们针对性地进行概念擦除。尽管概念擦除已在文本到图像和T2V模型的静态概念上得到广泛研究,但运动概念的擦除在很大程度上仍未探索。我们对视频扩散Transformer(DiTs)中的运动概念擦除进行了系统性研究。通过因果干预,我们表明文本条件注意力携带特定于概念的运动信息并支持选择性干预,而扰动时间位置编码则会同时抑制目标和非目标动态。我们进一步发现,将ESD(一种代表性的权重级图像擦除方法)直接适配到视频DiT上,会产生有限且不均匀的运动抑制:降低其擦除训练损失本身并不能从条件预测与无条件预测之间的差异中移除概念信号,而分类器自由引导(CFG)随后会在每个去噪步骤中放大该差异。基于这些发现,我们推导出运动概念擦除的三个要求:概念特异性、空间选择性和时间自然性。每个要求决定了MUTE(文本到视频生成中的运动概念遗忘)的一个组成部分:在每个去噪步骤中,MUTE通过令牌中和提取概念方向,从该方向的内在结构推导出空间门,并在应用CFG之前从速度输出中减去所得的修正。MUTE无需训练,且无需修改权重。在20个运动概念上的实验表明,MUTE在Wan2.1-T2V上优于代表性的提示级、权重级和推理时基线,且相同的公式可迁移到CogVideoX,支持其在不同T2V注意力架构中的适用性。

英文摘要

Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that motivate targeted concept erasure. Although concept erasure has been extensively studied for static concepts in text-to-image and T2V models, erasing motion concepts remains largely unexplored. We present a systematic study of motion concept erasure in video Diffusion Transformers (DiTs). Through causal interventions, we show that text-conditioning attention carries concept-specific motion information and supports selective intervention, whereas perturbing temporal positional encoding suppresses both target and non-target dynamics. We further find that directly adapting ESD, a representative weight-level image erasure method, to a video DiT yields modest and uneven motion suppression: reducing its erasure training loss does not by itself remove the concept signal from the difference between the conditional and unconditional predictions, which classifier-free guidance (CFG) then scales at every denoising step. From these findings, we derive three requirements for motion concept erasure: concept specificity, spatial selectivity, and temporal naturalness. Each determines one component of MUTE (Motion concept Unlearning in Text-to-video gEneration): at each denoising step, MUTE extracts a concept direction through token neutralization, derives a spatial gate from the direction's intrinsic structure, and subtracts the resulting correction from the velocity output before CFG is applied. MUTE is training-free and requires no weight modification. Experiments on 20 motion concepts show that MUTE outperforms representative prompt-level, weight-level, and inference-time baselines on Wan2.1-T2V, and the same formulation transfers to CogVideoX, supporting its applicability across distinct T2V attention architectures.

CommentsAccepted to ACM MM 2026. Dr. Chi Zhang is the corresponding author

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑