发表机构
University of Queensland; Adelaide University; Beijing Technology and Business University; City University of Hong Kong(昆士兰大学; 阿德莱德大学; 北京工商大学; 香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Caliber是针对返回分数API的跨架构模型提取防御方案,通过添加高斯噪声并校准噪声-效用关系,可控制替代模型性能,在30余组模型-数据集组合上误差低且防御效果符合预期。
AI 中文摘要
我们提出了Caliber,这是一种针对模型提取的输出扰动防御机制,它将噪声选择问题表述为一个校准问题:防御机制会降低用于训练替代模型的监督信号的程度,以及恢复干净logits的可证明的每输入查询成本。为了防御利用返回分数进行知识蒸馏的攻击者,Caliber向内部logits添加独立同分布的高斯噪声。我们确立了扰动后预测的两个性质:单调一致性退化,当干净logits具有唯一最大值时,与干净预测的一致性随噪声规模严格降低,因此(1/K,1)中的每个目标对应一个唯一的正规模;任务精度由可计算的上下包络线界定。每输入恢复成本,我们推导了固定输入下恢复干净logits所需重复查询的闭式极小极大下界。Caliber通过前两个logit差值的平方中位数对噪声方差进行归一化,并将得到的噪声-效用关系拟合为逻辑曲线,可按每个模型单独拟合或在任务内共享。在超过30个模型-数据集组合上,按模型校准实现了0.6-1.4%的平均绝对相对误差。端到端实验表明,替代模型性能通常与配置的退化程度一致,而固定输入平均则遵循预期的方差降低规律。
英文摘要
We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits. To defend against an attacker that uses returned scores for knowledge distillation, Caliber adds independent and identically distributed Gaussian noise to the internal logits. We establish two properties of the resulting perturbed predictions. Monotone agreement degradation: When the clean logits have a unique maximizer, agreement with the clean prediction decreases strictly with the noise scale, so every target in $(1/K,1)$ corresponds to a unique positive scale; task accuracy is bounded by computable lower and upper envelopes. Per-input recovery cost: We derive a closed-form minimax lower bound on the repeated queries needed to recover the clean logits for a fixed input. Caliber normalizes noise variance by the squared median top-two logit margin and fits the resulting noise-utility relationship with a logistic curve, either per model or shared within a task. Across more than thirty model-dataset combinations, per-model calibration achieves mean absolute relative errors of 0.6-1.4%. End-to-end experiments show that surrogate performance generally tracks the configured degradation, while fixed-input averaging follows the expected variance reduction.
Comments20 pages, 4 figures, 4 tables. Accepted at ProvSec 2026