COVER:基于生成式视频先验的编解码器鲁棒视频水印
COVER: Codec-Robust Video Watermarking with Generative Video Priors
浏览论文内容
中文总结 AI 辅助
COVER提出首个以编解码器压缩为设计目标的视频水印方法,通过生成式视频先验的潜在空间嵌入与可微编解码器替代库训练,在12种设置下平均比特准确率达93.72%,显著超越现有方法。
中文摘要 AI 辅助
视频水印支撑着生成媒体的版权保护与来源追溯,然而几乎每一段视频在存储或分享之前都会经过编解码器压缩。编解码器恰恰丢弃了大多数水印所依赖的感知冗余成分,因此即使标记视频在压缩前看起来完美无瑕,载荷也常常丢失。现有方法未能解决这一问题,因为它们将压缩视为通用失真列表中的一项,而真实的编解码器不可微,无法进入基于梯度的训练。我们提出COVER,这是首个以编解码器压缩为设计目标的 learned 视频水印方法,它通过将载荷嵌入冻结的生成式视频自编码器的潜在空间,并通过将接收到的视频重新编码到同一潜在空间来恢复载荷,从而在压缩后依然存活。为使编解码器鲁棒性可训练,我们构建了一个可微的编解码器替代库,模拟实际压缩的主要退化模式,并通过三条共享恢复路径在保真度目标下训练嵌入器和潜在解码器,该目标约束像素域和频域的残差。在4种编解码器、12种设置下,COVER达到93.72%的平均比特准确率,在12项中11项排名第一,比最强先前方法提升2.68个百分点,并将最差工作点从68.90%提升至73.72%,同时每个标记视频在视觉上保持与生成它的源片段接近。
英文摘要
Video watermarking underpins copyright protection and provenance for generated media, yet almost every video is compressed by a codec before it is stored or shared. A codec discards precisely the perceptually redundant components that most watermarks rely on, so the payload is often lost even when the marked video looked flawless beforehand. Existing methods leave this path open, since they treat compression as one entry in a generic list of distortions, while a real codec is not differentiable and cannot enter gradient-based training. We present COVER, the first learned video watermark built around codec compression as its design target, which survives that compression by embedding the payload in the latent space of a frozen generative video autoencoder and recovering it by re-encoding the received video into that same latent space. To make codec robustness trainable, we build a differentiable codec surrogate bank that simulates the dominant degradation modes of practical compression, and we train the embedder and the latent decoder through three shared recovery paths under a fidelity objective that constrains the residual in the pixel and frequency domains. Across four codecs at 12 settings, COVER attains 93.72% average bit accuracy, ranks first on 11 of the 12, improves the strongest prior method by 2.68 points, and lifts the worst operating point from 68.90% to 73.72% while each marked video stays visually close to the source clip that produced it.
发表机构
- National University of Singapore(新加坡国立大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。