深度神经网络中统计上不可检测的后门
Statistically Undetectable Backdoors in Deep Neural Networks
浏览论文内容
中文总结 AI 辅助
研究对抗模型训练者在深度前馈神经网络中植入统计上不可检测的后门,此后门可提供基于不变性的对抗样本,揭示了模型训练者与使用者间的权力不对称,且无后门时多项式时间内无法生成此类样本。
中文摘要 AI 辅助
我们展示了对抗模型训练者如何在一大类深度前馈神经网络中植入后门。这些后门在白盒设置下在统计上是不可检测的,即带后门和正常训练的模型在全变差距离上接近,即使给出模型的完整描述(如所有权重)。后门能为每个输入提供基于不变性的对抗样本,将不同输入映射到异常接近的输出。然而,没有后门,在多项式时间内(在标准密码学假设下)生成此类对抗样本是不可能的。我们的理论和初步实证结果证明了模型训练者和模型使用者之间存在根本的权力不对称。
英文摘要
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.
发表机构
- University of Ottawa(渥太华大学)
- Bocconi University(博科尼大学)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。