arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30980cs.CRcs.CV

FeatMark:基于扩散模型的特征级水印防护,抵御模仿攻击

FeatMark: Feature-level Watermark Protection against Mimicry Attacks with Diffusion Models

  • The Hong Kong Polytechnic University(香港理工大学)
  • CSIRO’s Data61(澳大利亚联邦科学与工业研究组织Data61研究所)
  • Macquarie University(麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

Haoyang Li, Ruoxi Sun, Qingqing Ye, Benjamin Zi Hao Zhao, Yaxin Xiao, Jason Xue, Haibo Hu

AI总结:

FeatMark提出特征级水印框架,通过场景一致的语义微特征替代像素扰动,抵御扩散模型模仿攻击,实验证明其鲁棒性高,可扩展至视频场景。

AI中文摘要:

文本到图像的扩散模型使得数据高效的“模仿”攻击成为可能,在这种攻击中,对手在少量公开照片上微调模型,以合成目标个体的令人信服的伪造品。一种常见的对策是嵌入不可感知的低能量水印,然而近期研究表明这些签名是脆弱的:适度的后处理或轻量级的对抗性扰动就能轻易抑制检测,这暴露了不可感知性与鲁棒性之间的根本矛盾。我们提出了FeatMark,一个水印框架,它将从像素级、能量受限的扰动转向不显眼的语义特征:小而场景一致的微特征,这些特征对人类来说保持自然,同时提供更强、机器可验证的来源信号。FeatMark构建了特定领域的特征库,将每个水印编码为紧凑的概念程序,将开放词汇的语义线索与可靠的编辑区域和指令模板配对。然后,它自动选择既可行又可执行的特征,并通过模块化的、掩码引导的概念编辑注入这些特征,产生高度局部化、场景一致的微编辑,这些编辑难以被察觉。我们在VGGFace2、CelebA-HQ和WikiArt上进行了广泛的实验,评估了10种强水印移除/净化攻击(包括再生式净化)和几种针对FeatMark定制的自适应攻击,以评估感知保真度、水印检测准确性和鲁棒性。我们进一步展示了FeatMark对视频模仿攻击的可扩展性。结果表明,FeatMark几乎不可渗透,能够抵御所有评估的攻击,且比特准确性和保真度退化可忽略不计。

英文摘要:

Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We introduce FeatMark, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro-features that remain natural to humans while providing a stronger, machine-verifiable provenance signal. FeatMark builds domain-specific feature banks that encode each watermark as a compact concept program, pairing open-vocabulary semantic cues with reliable edit regions and instruction templates. It then automatically selects features that are both feasible and executable and injects them through modular, mask-guided concept editing, yielding highly localized, scene-consistent micro-edits that are difficult to perceive. We conduct extensive experiments across VGGFace2, CelebA-HQ, and WikiArt, evaluating against 10 strong watermark removal/purification attacks (including regeneration-style purification) and several bespoke adaptive attacks tailored to FeatMark, to assess perceptual fidelity, watermark detection accuracy, and robustness. We further demonstrate FeatMark's extensibility to video mimicry attacks. The results show FeatMark remains virtually impervious, withstanding all evaluated attacks with negligible bit-accuracy and fidelity degradation.

补充信息

↑