AI 中文总结
针对现有水印攻击难以捕捉全局水印长程依赖、攻击效果与保真度难平衡的问题,提出SPFM-Net,通过高比率掩码、掩码自编码器、多尺度残差频率特征交互及GSFM单元实现水印抑制,在多类水印方案上取得良好平衡。
AI 中文摘要
现有的水印攻击通常依赖于预定义的信号处理操作或局部约束的修复网络,难以捕捉全局分布的水印信号的长程依赖关系,且在去除效果与视觉保真度之间难以取得良好平衡。本文提出SPFM-Net,一种语义先验引导且频率约束的Mamba框架,用于不可见水印攻击。SPFM-Net首先采用高比率掩码破坏不可见水印信号的空间连贯性,随后利用部分微调的预训练掩码自编码器从稀疏观测中重建语义一致的图像,同时抑制与水印相关的信息。多尺度残差频率特征交互模块随后聚合多个感受野中与水印相关的残差特征,同时自适应抑制与水印无关区域的响应。为进一步捕捉全局分布水印信号的长程依赖关系,引入轻量型基于Mamba的全局状态空间特征建模(GSFM)单元,以分离水印相关特征与自然图像内容并抑制剩余水印痕迹。此外,SPFM-Net采用多级目标函数优化,联合施加空间、频率和边缘域约束,在实现有效水印抑制的同时保留感知质量。在代表性的空间域、变换域、基于正交矩及深度学习的水印方案上开展的大量实验表明,SPFM-Net在水印攻击效果与感知保真度之间取得了良好的平衡。
英文摘要
Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack. SPFM-Net first employs high-ratio masking to disrupt the spatial coherence of invisible watermark signals, and then utilizes a partially fine-tuned pretrained Masked Autoencoder to reconstruct semantically consistent image from sparse observations while suppressing watermark-related information. A Multi-scale Residual Frequency Feature Interaction module subsequently aggregates watermark-related residual features across multiple receptive fields, while adaptively suppressing responses from watermark-irrelevant regions. To further capture the long-range dependencies of globally distributed watermark signals, a lightweight Mamba-based Global State-space Feature Modeling (GSFM) unit is introduced to separate watermark-related features from natural image content and suppress the remaining watermark traces. In addition, SPFM-Net is optimized using a multi-level objective that jointly imposes spatial-, frequency-, and edge-domain constraints, enabling effective watermark suppression while preserving perceptual quality. Extensive experiments on representative spatial-domain, transform-domain, orthogonal moment-based, and deep learning-based watermarking schemes demonstrate that SPFM-Net achieves a favorable trade-off between watermark attack effectiveness and perceptual fidelity.