FOMO:在选择性视频遗忘中遗忘概念,不错过场景
FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning
- Jagiellonian University(雅盖隆大学)
- IDEAS Research Institute(IDEAS 研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视频生成模型中的有害内容,提出首个基于训练的选择性视频遗忘方法FOMO,在移除目标概念(含运动概念)的同时优先保留原始场景,实现最佳权衡。
AI中文摘要:
生成式视频模型的快速发展使得合成越来越逼真且时间连贯的视频成为可能,同时也引发了对有害内容生成的担忧。训练过程中对大规模网络数据集的依赖不可避免地使这些模型接触到不良材料,使得概念遗忘成为一项必要的缓解措施。现有方法主要针对静态视觉概念,如物体、身份或不安全外观,在很大程度上忽视了运动遗忘。此外,这些方法往往很少关注对周围场景的保留。因此,成功的概念移除可能会无意中改变背景、构图或整体视频动态。我们认为,有效的遗忘理想情况下应只改变目标内容,同时最小化对剩余场景的不必要改变。在这项工作中,我们引入了FOMO,据我们所知,这是第一种基于训练的选择性视频遗忘方法,直接将对原始场景的保留作为优先事项。我们围绕两个互补的目标来构建遗忘:改变什么和保留什么。我们的方法定位与概念相关的表示并对其进行修改,而保留机制则在不需辅助数据的情况下维持非目标场景信息。除了简单地擦除不需要的概念外,FOMO还明确地将生成引导到指定的安全替代方案。我们进一步将此公式扩展到运动遗忘,其中概念由时间行为而非固定空间区域定义。我们的解决方案在不安全内容、物体和运动概念上实现了有效的遗忘,同时在概念移除和场景保留之间取得了最佳权衡。代码:此 https URL 项目页面此 https URL
英文摘要:
The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesirable material, making concept unlearning an essential mitigation. Existing methods mainly target static visual concepts, such as objects, identities, or unsafe appearance, largely overlooking motion unlearning. Furthermore, these approaches often pay little attention to preserving the surrounding scene. As a result, successful concept removal may unintentionally alter the background, composition, or overall video dynamics. We argue that effective unlearning should ideally change only what is targeted, while minimizing unnecessary changes to the remaining scene. In this work, we introduce FOMO, to the best of our knowledge the first training-based selective video unlearning method that directly treats preservation of the original scene as a priority. We formulate unlearning around two complementary objectives: what to change and what to preserve. Our method localizes concept-related representations and modifies them, while the preservation mechanism maintains non-target scene information without requiring auxiliary data. Beyond simply erasing unwanted concepts, FOMO explicitly redirects the generation toward a specified safe alternative. We further extend this formulation to motion unlearning, where the concept is defined by temporal behavior rather than a fixed spatial region. Our solution achieves effective unlearning across unsafe content, object, and motion concepts, while achieving the best trade-off between concept removal and scene preservation. Code: https://github.com/gmum/FOMO Project Page https://gmum.github.io/FOMO