发表机构
University of Fribourg; University of Nis; University MB(弗里堡大学; 尼什大学; MB大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出双锁定方法,结合稀疏QIM水印与自适应索引置换,实现模型保护、恢复和所有权验证,实验显示锁定后准确率大幅下降,正确密钥可完全恢复。
AI 中文摘要
我们提出了一种用于保护已训练神经网络的双锁定方法,该方法将密钥驱动的索引置换与基于稀疏量化索引调制(Sparse Quantization Index Modulation, QIM)的PIN水印相结合。通过独立地对自适应选择的索引向量的每一行应用均匀随机置换,引入了密码学随机性。随后,通过调制偏置系数的量化值,将鲁棒的盲二进制水印嵌入其中,从而将网络与用户定义的个人识别码(PIN)绑定。在没有正确密钥的情况下,网络保留其架构,但由于内部表示被破坏而功能受损。逆置换可完全恢复原始模型精度,同时嵌入的水印保持不可感知,并能够实现对密钥关联和模型作者身份的盲验证。为提高锁定效果和可恢复性,一种自适应密钥选择策略将高幅值权重重新分配到低敏感度位置,反之亦然,从而在锁定状态下增加性能退化,同时保持完全恢复能力。在MNIST、CIFAR-10/100和ImageNet-1K上,使用全连接网络、ResNet CNN和Transformer架构进行的实验表明,锁定使准确率降至10%以下,对于CNN甚至低于0.5%,而正确密钥可完全恢复性能。水印未引入可测量的精度退化,并能可靠地验证所有权。对CNN和Transformer嵌入分布的分析进一步表明,其具有识别训练不足或次优设计模型的潜在诊断价值。因此,所提出的方法同时提供了模型保护、恢复和所有权验证。
英文摘要
We present a dual-locking method for securing trained neural networks that combines key-driven index permutation with PIN-based watermarking based on Sparse Quantization Index Modulation (QIM). Cryptographic randomness is introduced by independently applying a uniform random permutation to each row of adaptively selected index vectors. A robust blind binary watermark is then embedded into the bias coefficients by modulating their quantized values, binding the network to a user-defined Personal Identification Number (PIN). Without the correct key, the network retains its architecture but becomes functionally impaired due to disrupted internal representations. Inverse permutation fully restores the original model accuracy, while the embedded watermark remains imperceptible and enables blind verification of key association and model authorship. To improve both locking effectiveness and recoverability, an adaptive key selection strategy redistributes high-magnitude weights to low-sensitivity positions and vice versa, increasing degradation in the locked state while preserving full recovery. Experiments on MNIST, CIFAR-10/100, and ImageNet-1K using fully connected networks, ResNet CNNs, and transformer architectures show that locking reduces accuracy below 10\%, and even below 0.5\% for CNNs, while the correct key fully restores performance. The watermark introduces no measurable accuracy degradation and reliably authenticates ownership. Analysis of embedding distributions across CNNs and transformers further indicates potential diagnostic value for identifying undertrained or suboptimally designed models. The proposed approach therefore provides simultaneous model protection, recovery, and ownership verification.
Comments14 pages, 5 figures, 7 tables, IEEE TAI
Journal refIEEE Transactions on Artificial Intelligence, vol. 7, no. 6, pp. 3259-3272, June 2026