对抗稀疏教师:使用对抗样本防御基于蒸馏的模型窃取攻击
Adversarial Sparse Teacher: Defense Against Distillation-Based Model Stealing Attacks Using Adversarial Examples
- Hacettepe University(哈杰泰佩大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出对抗稀疏教师(AST)方法,利用对抗样本和EPD损失函数增加输出熵,有效防御基于蒸馏的模型窃取攻击,并在CIFAR数据集上保持高精度。
AI中文摘要:
我们引入了对抗稀疏教师(AST),一种针对基于蒸馏的模型窃取攻击的稳健防御方法。我们的方法使用对抗样本来训练教师模型,以产生稀疏的逻辑响应并增加输出分布的熵。通常,模型会在其输出中产生一个对应于其预测的峰值。通过利用对抗样本,AST修改教师模型的原始响应,在输出中嵌入少量改变的逻辑值,同时保持主要响应略高。同时,所有剩余的逻辑值被提升以进一步增加输出分布的熵。所有这些复杂的操作都是通过我们提出的指数预测散度(EPD)损失函数的优化函数来执行的。与传统KL散度相比,EPD允许我们维持更高的熵水平,从而有效迷惑攻击者。在CIFAR-10和CIFAR-100数据集上的实验表明,AST优于最先进的方法,在保持高精度的同时提供了对模型窃取的有效防御。源代码将很快在此公开。
英文摘要:
We introduce Adversarial Sparse Teacher (AST), a robust defense method against distillation-based model stealing attacks. Our approach trains a teacher model using adversarial examples to produce sparse logit responses and increase the entropy of the output distribution. Typically, a model generates a peak in its output corresponding to its prediction. By leveraging adversarial examples, AST modifies the teacher model's original response, embedding a few altered logits into the output while keeping the primary response slightly higher. Concurrently, all remaining logits are elevated to further increase the output distribution's entropy. All these complex manipulations are performed using an optimization function with our proposed Exponential Predictive Divergence (EPD) loss function. EPD allows us to maintain higher entropy levels compared to traditional KL divergence, effectively confusing attackers. Experiments on CIFAR-10 and CIFAR-100 datasets demonstrate that AST outperforms state-of-the-art methods, providing effective defense against model stealing while preserving high accuracy. The source codes will be made publicly available here soon.