发表机构
Macau University of Science and Technology; FoShan University; Macao Polytechnic University(澳门科技大学; 佛山大学; 澳门理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对深度学习版权保护面临的非法模型训练与数据泄露威胁,该文提出一种可逆不可学习样本机制,通过生成诱导模型学习无关特征的扰动与双水印提取策略,在三类数据集上实现全面版权保护。
AI 中文摘要
深度学习的显著进步得益于大规模数据集的使用,这凸显了版权保护的关键重要性。向样本添加精心设计的扰动使其不可学习,已成为保护数据版权的重要方法。现有的不可学习样本生成方法忽略了数据泄露风险,这可能威胁数据所有权。因此,深度学习中的版权保护面临两大主要威胁:非法模型训练和恶意数据泄露。我们研究发现,现有可用性攻击与水印技术的简单组合无法解决这两种威胁,因为它们存在负面交互效应。因此,本文针对上述安全问题提出了一种新颖的版权保护机制。考虑到防止未经授权的模型训练需要不可学习扰动具有强大的泛化性,我们生成扰动以诱导模型学习输入图像的不相关特征,其原理是最小化模型输入与输出之间的互信息。另一方面,为消除不可学习扰动对水印提取的副作用,我们设计了一种双提取策略,使用两个不同的水印提取器。在ImageNet、CIFAR10和Pets图像数据集上的大量实验表明,所提方法可为图像提供全面的版权保护。代码可在该https链接获取。
英文摘要
Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data copyright. Existing methods for creating unlearnable examples overlook the risk of data leakage, which can threaten data ownership. Thus, copyright protection in deep learning faces two main threats: illegal model training and malicious data leakage. We investigate that these two threats cannot be solved by straightforwardly combining existing availability attacks and watermarking techniques as their negative interaction effects. Therefore, in this paper, we propose a novel copyright protection mechanism for the aforementioned security concerns. Considering that the prevention of unauthorized model training requires powerful generalizability of unlearnable perturbations, we generate perturbations to induce the model to learn uncorrelated features of input images. It works by minimizing the mutual information of the input and output of the model. On the other hand, to eliminate the side impact of unlearnable perturbations on the watermark extraction, we design a dual extraction strategy by using two distinct watermark extractors. Extensive experiments on the image datasets {ImageNet, CIFAR10, and Pets} show that our proposed method could provide comprehensive copyright protection to images. The code is available at {https://github.com/Yeah21/ReversibleUnlearnableExamples}.