发表机构
Rochester Institute of Technology(罗彻斯特理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对知识蒸馏中对手复制专有分类器的问题,提出ADS-C方法,通过在闭式的每个输入边际预算下构成扰动,保留服务的top-1预测,使防御后教师模型准确率不变,且效用成本为零,有效抵御对手蒸馏,降低学生模型性能损失。
AI 中文摘要
知识蒸馏使对手能够通过查询专有分类器的预测接口并根据返回的概率向量训练替代模型来复制它。针对大语言模型提出的反蒸馏采样通过对服务分布进行依赖输入、梯度导向的扰动来应对这种威胁,但其在分类中的转移尚未得到研究。将这种防御方法应用于分类时,我们发现其行为受教师模型每个输入的置信度边际分布的支配。由于训练良好的分类器存在严重的过度自信问题,直接转移存在一个惰性窗口:低于一个可闭式预测的阈值时,它对攻击者和防御者都没有影响;超过该阈值,防御会发生相变,并且比攻击者的学生模型更快地降低教师模型的性能。温度软化以闭式重新调整了这种转变,并且每个温度配置都处于相同的不利权衡曲线上。我们的方法ADS-C在闭式的每个输入边际预算下构成扰动,可证明保留每个服务的top-1预测,因此防御后的教师模型的准确率与未防御的教师模型相同。在此保证下,蒸馏后的学生模型在CIFAR-100上仍损失17.4个百分点,在CIFAR-10上损失29.6个百分点,在Tiny-ImageNet上损失13.3个百分点;与未修改的防御方法匹配这种性能下降会使教师模型的准确率损失27.5、32.9和22.2个百分点。由于服务标签不变,硬标签攻击者没有收获,而防御后的软输出训练的学生模型比该下限低29.7个百分点:蒸馏服务概率的动机不仅被消除,而且被逆转。据我们所知,ADS-C是第一种效用成本恰好为零的分类反蒸馏防御方法。
英文摘要
Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors. Antidistillation sampling, proposed for large language models, counters this threat with an input-dependent, gradient-directed perturbation of the served distribution; its transfer to classification has not been studied. Adapting the defense to classification, we show its behavior is governed by the distribution of the teacher's per-input confidence margins. Because well-trained classifiers are severely overconfident, the direct transfer exhibits an inert window: below a closed-form-predictable threshold, it affects neither attacker nor defender; beyond it, the defense undergoes a phase transition and degrades the teacher faster than the attacker's student. Temperature softening rescales the transition in closed form, and every temperature configuration lies on the same unfavorable trade-off curve. Our method, ADS-C, composes the perturbation under a closed-form, per-input margin budget that provably preserves every served top-1 prediction, so the defended teacher's accuracy equals the undefended teacher's identically. Under this guarantee the distilled student still loses 17.4 percentage points on CIFAR-100, 29.6 on CIFAR-10, and 13.3 on Tiny-ImageNet; matching this degradation with the unmodified defense costs 27.5, 32.9, and 22.2 points of teacher accuracy. Because served labels are unchanged, a hard-label attacker gains nothing, while the defended soft output trains a student up to 29.7 points below that floor: the incentive to distill served probabilities is not merely removed but reversed. To our knowledge, ADS-C is the first antidistillation defense for classification whose utility cost is exactly zero.
Comments20 pages, 28 figures