arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03174cs.CRcs.AI

面向生成式AI模型的基于属性的不可检测水印

Attribute-based Undetectable Watermarking for Generative AI Models

Miryam Mi-Ying Huang, Chung-Wei Lee, Max Raffel, Er-Cheng Tang

首次发表
浏览论文内容

中文总结 AI 辅助

针对生成式AI水印中检测密钥滥用的问题,提出首个基于属性的生成式AI水印方案,实现策略控制的细粒度检测,经原型验证兼具有效性与实用性。

中文摘要 AI 辅助

生成式AI系统生成的内容来源难以验证的问题日益突出,这推动了用于识别模型生成输出的水印技术的发展。现有的密码学水印方法提供了强不可检测性保证:在没有检测密钥的情况下,带水印的输出与未加水印的输出在计算上不可区分。然而,这些方法未解决如何安全委托检测能力这一关键部署挑战。若检测密钥不受限制,恶意检测器可能会将检测密钥用于超出预期范围的用途,从而实现水印清除、范围滥用和用户画像。为缓解这一安全问题,我们提出了据我们所知首个面向生成式AI模型的基于属性的水印,提供细粒度、策略控制的水印检测。在我们的方法中,每个生成的输出都与属性关联,每个检测密钥都受潜在属性上的策略约束。检测密钥仅可用于检测属性满足对应策略的带水印输出,而落在策略之外的带水印输出仍与未加水印的输出在计算上不可区分。我们构建了此类基于属性的水印方案,并形式化了其安全属性,包括一致性、对有界腐败的自适应鲁棒性、不可检测性和正确性,同时在标准密码假设下提供了安全证明。我们的构建将受限伪随机函数、伪随机纠错码和随机性恢复过程与生成式AI模型相结合。最后,我们实现了原型并进行了实证评估,证明基于属性的水印既有效又实用。

英文摘要

Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identifying model-generated outputs. Existing cryptographic watermarking methods provide strong undetectability guarantees: without a detection key, watermarked outputs are computationally indistinguishable from unwatermarked ones. However, these approaches do not address the crucial deployment challenge of how to safely delegate detection capabilities. With an unrestricted detection key, a malicious detector may use the detection key beyond its intended scope, enabling watermark sanitization, scope abuse, and user profiling. To mitigate this safety concern, we introduce, to the best of our knowledge, the first \emph{attribute-based watermarking} for generative AI models, providing fine-grained, policy-controlled watermark detection. In our approach, each generated output is associated with attributes, and each detection key is \emph{constrained by a policy} on potential attributes. A detection key can only be used to detect watermarked outputs whose attributes satisfy the corresponding policy, while watermarked outputs that fall outside the policy remain computationally indistinguishable from unwatermarked ones. We construct such an attribute-based watermarking scheme and formalize its security properties, including consistency, adaptive robustness to bounded corruptions, undetectability, and soundness, along with a security proof under standard cryptographic assumptions. Our construction integrates constrained pseudorandom functions, pseudorandom error-correcting codes, and randomness recovery procedures with generative AI models. Finally, we implement a prototype and an empirical evaluation, demonstrating that attribute-based watermarking is both effective and practical.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Southern California(南加州大学)
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑