为什么对神经网络进行后门攻击如此容易?
Why Backdooring Neural Networks is so Easy?
浏览论文内容
中文总结 AI 辅助
本文通过二次神经元的闭式分析,揭示特征学习使神经网络更易受后门攻击,其所需触发器强度随投毒比例按α∝π^{-1/4}缩放,远低于惰性学习的π^{-1/2},并指出线性启发式安全审计会低估此风险。
中文摘要 AI 辅助
保护现代AI系统免受后门攻击仍然是一个开放性的挑战,需要从根本上对攻击者的预算进行原则性估计——即构造成功且隐蔽的攻击所需的投毒比例π和触发器强度α。受近期经验证据的启发,该证据表明,即使干净数据集不断增长,对大型语言模型进行投毒也可能需要近乎恒定数量的恶意样本,我们对在投毒高斯混合上训练的二次神经元进行了精确的闭式分析。我们表明,也许与直觉相反,使神经网络强大的相同特征学习动态也可能使其更容易受到后门攻击。具体来说,在干净准确率保持一阶O(π)的情况下,我们证明了惰性学习对成功攻击施加了逆平方根缩放α∝π^{-1/2},而特征学习则诱导出一个二次检测器,其损失边际按O(α^4)缩放,将攻击预算改进为α∝π^{-1/4}。因此,非线性特征学习在小投毒比例下大幅降低了所需的触发器强度,从而在某种意义上使特征学习者更容易受到后门攻击。这些结果提供了一种与大规模经验观察一致的理论机制,并表明基于线性启发式的安全审计会系统性地低估广泛采用的特征学习机制中的后门漏洞。
英文摘要
Securing modern AI systems against backdoor attacks remains an open challenge and requires fundamentally principled estimates of the adversary's budget -- the poison fraction $π$ and trigger strength $α$ needed to construct successful yet stealthy attacks. Motivated by recent empirical evidence that poisoning large language models can require a nearly constant number of malicious samples even as clean datasets grow, we derive an exact closed-form analysis of a quadratic neuron trained on a poisoned Gaussian mixture. We show, perhaps counterintuitively, that the same feature-learning dynamics that make neural networks powerful can also make them more vulnerable to backdoors. Specifically, with clean accuracy preserved to first order, $O(π)$, we demonstrate that lazy learning imposes the inverse-square-root scaling $α\propto π^{-1/2}$ for a successful attack, while feature learning induces a quadratic detector whose loss margin scales as $O(α^4)$, improving the attack budget to $α\propto π^{-1/4}$. Consequently, nonlinear feature learning substantially reduces the trigger strength required at small poison fractions, thereby in a sense making feature learners more backdoor vulnerable. These results provide a theoretical mechanism consistent with large-scale empirical observations and demonstrate that security audits based on linear heuristics can systematically underestimate backdoor vulnerability in the widely adopted feature-learning regimes.
发表机构
- Université Paris-Saclay, CEA LIST(巴黎-萨克雷大学,法国原子能委员会信息与系统技术实验室)
- AI Research Center, Technology Innovation Institute (TII)(人工智能研究中心,技术创新研究所)
机构由 AI 辅助整理,请以论文原文为准。