arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeepFreqMark:面向潜在扩散模型的、结合球形攻击模拟的端到端可学习频域水印

DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models

Chen-Hsiu Huang, Mario Köppen, Ja-Ling Wu

arXiv 2608.08999首次发表:更新:

发表机构

National Taiwan University; Kyushu Institute of Technology(台湾大学; 九州工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对潜在扩散模型生成图像的版权与虚假信息问题,提出DeepFreqMark频域水印框架,结合球形攻击模拟规避计算瓶颈,其误码率显著低于基线,消息容量可达256位。

AI 中文摘要

由潜在扩散模型(LDM)生成的AI图像的大量涌现,引发了关于版权侵权和虚假信息的关键担忧。尽管现有的频域水印方法会在生成前的初始潜在噪声中嵌入手工制作的几何图案,但它们存在容量有限、图案设计僵化的问题。我们提出DeepFreqMark,这是一个端到端可学习的频域水印框架,用神经消息编码器和解码器取代手动图案工程。为了规避训练期间去噪扩散隐式模型(DDIM)逆过程造成的计算瓶颈,我们引入了基于球面线性插值(Slerp)的攻击模拟,该方法直接在噪声潜在上操作,同时严格保留高斯方差。大量实验表明,DeepFreqMark在现实攻击下的误码率(BER)显著低于基线方法,且可扩展至256位的消息容量。我们的源代码可在该https URL获取。

英文摘要

The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infringement and misinformation. Although existing frequency-domain watermarking methods embed handcrafted geometric patterns into the initial latent noise prior to generation, they suffer from limited capacity and rigid pattern designs. We propose DeepFreqMark, an end-to-end learnable frequency-domain watermarking framework that replaces manual pattern engineering with a neural message encoder and decoder. To circumvent the computational bottleneck caused by Denoising Diffusion Implicit Model (DDIM) inversion during training, we introduce a Spherical Linear Interpolation (Slerp)-based attack simulation. This approach operates directly on the noise latent while strictly preserving the Gaussian variance. Extensive experiments demonstrate that DeepFreqMark achieves significantly lower Bit Error Rates (BER) than baseline methods under real-world attacks and scales to 256 bits message capacity. Our source code is available at https://github.com/chenhsiu48/DeepFreqMark.

Commentsaccepted by APSIPA ASC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑