发表机构
Tsinghua University; Shandong University; Wuhan University; Zhongguancun Laboratory; Shandong Institute of Blockchain; National Financial Cryptography Research Center(清华大学; 山东大学; 武汉大学; 中关村实验室; 山东省区块链研究院; 国家金融密码研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出GhostVAE方法,通过在VAE编码器植入隐秘后门,可在良性图像上保持94.4%的平均水印真阳性率,同时实现94.6%的平均水印规避成功率,突破了现有语义水印的安全性。
AI 中文摘要
尽管语义水印被视为对潜在扩散模型(LDMs)生成图像的一种有前景的保护措施,但水印检测流程对神经网络的依赖引入了一个关键却未被充分探索的后门攻击面。为系统研究这一漏洞,我们提出GhostVAE,以在变分自编码器(VAE)的编码器中植入隐秘后门,从而可靠地规避水印检测。GhostVAE分两个阶段运行:首先,它通过功率谱正则化构建通用触发器,以提升触发器的鲁棒性;随后,通过参数对齐目标训练中毒的VAE编码器。我们在三种最先进的语义水印方案和三种广泛使用的LDMs上开展大量评估,结果显示,GhostVAE在良性图像上保持水印检测性能(平均真阳性率达94.4%),同时在触发器激活下实现高度有效的规避(平均攻击成功率达94.6%)。此外,我们对17种代表性防御措施进行全面分析,证明GhostVAE在输入空间、参数空间和潜在空间中仍保持隐秘性。我们的工作从根本上削弱了语义水印系统的可信度,并强调语义水印的安全部署需要端到端的安全考量,尤其是针对神经网络组件。
英文摘要
Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.
CommentsTo appear in USENIX Security 2026, August 12-14, 2026