arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14194cs.CVcs.LG

文本到视频模型的推理时概念抑制和以视频为中心的评估

Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models

发表机构中国科学技术大学
查看机构详情
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenxuan Chen, Wenjie Feng

首次发表
浏览论文内容

中文总结 AI 辅助

研究文本到视频模型中可控去除目标概念的问题,提出无训练推理时框架SIRUS,通过定位目标相关提示证据抑制目标表达,引入视频导向评估框架,在多个概念上实验表现出色,能在不同骨干上泛化。

中文摘要 AI 辅助

文本到视频(T2V)生成器能合成逼真且时间连贯的视频,但可控地从生成器中去除目标概念仍很困难。与文本到图像概念擦除不同,T2V的遗忘必须在保留非目标主体、动作、场景和时间结构的同时抑制可能跨帧持续的目标概念。我们提出了SIRUS,一个用于概念级T2V遗忘的无训练推理时框架。给定目标概念的文本别名,SIRUS定位与目标相关的提示证据并在采样期间抑制目标表达,而不更新文本编码器或去噪网络。我们还引入了一个面向视频的T2V遗忘评估框架,使用视频级失败标准、帧级残差统计、配对保留分析、基于VBench的质量诊断和部署开销测量,分别测量目标遗忘、非目标保留、视频质量、越狱鲁棒性和效率。在CogVideoX上的五个安全、对象和风格概念上,SIRUS平均遗忘成功率达到70.4%,平均帧命中率达到25.7%,而VideoEraser分别为44.4%/47.2%,同时将平均VBench质量下降从-0.043降低到-0.016,在所有评估的基线中实现了最强的遗忘-质量权衡。在Wan2.2上的迁移实验进一步表明SIRUS能在现代T2V骨干上进行泛化。

英文摘要

Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult. Unlike text-to-image concept erasure, T2V unlearning must suppress a target concept that may persist across frames while preserving non-target subjects, actions, scenes, and temporal structure. We propose \textbf{SIRUS}, a training-free inference-time framework for concept-level T2V unlearning. Given textual aliases of a target concept, SIRUS localizes target-related prompt evidence and suppresses target expression during sampling, without updating the text encoder or denoising network. We further introduce a video-oriented evaluation framework for T2V unlearning that separately measures target forgetting, non-target preservation, video quality, jailbreak robustness, and efficiency, using video-level failure criteria, frame-level residue statistics, paired preservation analysis, VBench-based quality diagnostics, and deployment overhead measurement. Across five safety, object, and style concepts on CogVideoX, SIRUS reaches 70.4\% average forgetting success and 25.7\% average frame hit, compared with 44.4\% / 47.2\% for VideoEraser, while reducing the average VBench quality drop from -0.043 to -0.016, yielding the strongest forgetting-quality trade-off among fully evaluated baselines. Transfer experiments on Wan2.2 further suggest that SIRUS generalizes across modern T2V backbones.

补充信息

↑