arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17747cs.CV

TINA+: 通过扩散一致无文本反演探测未学习的扩散模型中的残留视觉知识

TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

Qianlong Xiang, Miao Zhang, Kun Wang, Haoyu Zhang, Junhui Hou, Liqiang Nie

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出TINA+,一种扩散一致无文本反演攻击,可探测未学习的扩散模型中残留的视觉知识,实验表明现有概念擦除方法多切断文本-图像链接而非消除视觉知识。

中文摘要 AI 辅助

尽管文本到图像的扩散模型展现出卓越的生成能力,但概念擦除技术对于防止有害内容至关重要。现有的对抗性探针通过测试被擦除的概念是否仍能被恢复来评估这些方法。然而,现有的擦除和探针方法在很大程度上仍以文本为中心,侧重于文本到图像的映射是否被切断,而忽略了相应的视觉知识是否保留。为了从视觉角度研究这一问题,我们利用扩散反演来探测生成轨迹是否能重建被擦除概念的视觉实例。在空文本条件下,标准反演避开了文本路径,但放大了近似误差,阻碍了忠实轨迹的恢复。为应对这一挑战,我们引入了TINA+,一种配备基于优化的反演的扩散一致无文本反演攻击。我们还发现,无约束的扩散反演可能会发现虚假轨迹,甚至允许随机初始化的扩散模型重建目标概念,这类轨迹可能会错误地指示残留视觉知识。因此,TINA+引入了扩散一致轨迹正则化来抑制这种失效模式,通过惩罚远低于扩散预期边际能量演化的轨迹,TINA+在抑制虚假反演路径的同时保留了恢复被擦除概念的能力。在12种擦除方法、4项概念擦除任务和不同模型架构上进行的实验表明,TINA+能通过扩散一致视觉轨迹可靠地探测残留视觉知识,这些结果提供了更有力的证据,表明当前方法通常通过切断文本-图像链接来模糊概念,而非消除潜在的视觉知识。

英文摘要

Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely text-centric, focusing on whether the text-to-image mapping is severed while overlooking whether the corresponding visual knowledge remains. To investigate this question from a visual perspective, we leverage diffusion inversion to probe whether a generative trajectory can reconstruct visual instances of an erased concept. Under a null-text condition, standard inversion avoids the textual pathway but amplifies approximation errors, hindering faithful trajectory recovery. To address this challenge, we introduce TINA+, a diffusion-consistent Text-free INversion Attack equipped with optimization-based inversion. We also find that unconstrained diffusion inversion may discover spurious trajectories, even allowing a randomly initialized diffusion model to reconstruct the target concept. Such trajectories may falsely indicate residual visual knowledge. TINA+ therefore introduces Diffusion-Consistent Trajectory Regularization to suppress this failure mode. By penalizing trajectories that fall far below the expected marginal energy evolution of diffusion, TINA+ suppresses spurious inversion paths while preserving its ability to recover erased concepts. Experiments across twelve erasure methods, four concept-erasure tasks, and different model architectures demonstrate that TINA+ reliably probes residual visual knowledge through diffusion-consistent visual trajectories. These results provide stronger evidence that current methods often obscure concepts by severing text-image links rather than eliminating the underlying visual knowledge.

发表机构

  • School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)
  • City University of Hong Kong(香港城市大学)
  • Shenzhen Loop Area Institute(深圳河套学院)
  • National University of Singapore(新加坡国立大学)
  • Pengcheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑