arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39548cs.CVcs.AIcs.CR

学习正常扩散动力学用于文本到图像模型的后门防御

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

  • Geely(吉利)
  • Hunan University(湖南大学)
  • East China Normal University(华东师范大学)
  • University of Adelaide(阿德莱德大学)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • China University of Petroleum (East China)(中国石油大学(华东))

机构由 AI 辅助整理,请以论文原文为准。

Junjian Li, Xiaolong Liu, Peng Sun, Liantao Wu, Linghan Chen, Yudong Gao, Honglong Chen

AI总结:

本文提出NDDL框架,从转移动力学视角学习良性扩散轨迹的正常模式,通过预测偏差检测后门并定位触发器,实验验证其有效性与泛化性。

AI中文摘要:

后门攻击对文本到图像(T2I)扩散模型的安全部署构成严重威胁。现有防御方法通常从内部表示中的特定异常模式检测后门,这可能随着越来越多样的攻击机制的出现而限制其泛化能力。在本文中,我们从转移动力学的角度研究T2I扩散模型的后门防御。我们观察到良性扩散轨迹在交叉注意力、潜空间和噪声空间中表现出结构化的、依赖于时间步长的转移模式,而后门攻击往往会导致偏离这种正常演化。受这些观察的启发,我们提出了正常扩散动力学学习(NDDL),一种新颖的后门防御框架,仅利用良性样本学习扩散轨迹的正常转移动力学。NDDL构建紧凑的多空间轨迹表示,并训练一个时间步长条件动力学模型来预测扩散演化。在推理阶段,利用观测与预测转移之间的偏差来量化动力学不一致性,用于后门检测。NDDL还通过在无任何嵌入后门先验知识的情况下执行低语义词替换,实现触发器定位。针对多种后门攻击的大量实验证明了我们提出的NDDL的有效性和泛化能力。

英文摘要:

Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.

↑