AI 中文总结
该研究提出针对持续学习网络的新型攻击,定义学习阻断器与灾难性学习,提出六种攻击策略,经超4480次模拟验证可阻碍新知识获取并加剧旧知识丧失。
AI 中文摘要
持续学习(Continual Learning, CL)使深度学习模型能够从数据流中迭代学习,同时不会遗忘先前的知识。现有针对CL的对抗性研究主要旨在重新引发灾难性遗忘,攻击其稳定性并降低可用性。我们发现了一种新型安全漏洞:攻击者操纵的数据会降低当前或后续迭代的可学习性,我们将此类操纵称为学习阻断器(learning blockers),因为它们攻击CL算法的可塑性。这类操纵尤其有害,因为在当前迭代的训练过程中难以检测,它们可针对模型尚未遇到数据的迭代。当学习阻断器还引发灾难性遗忘时,由此产生的整体退化就是我们所说的灾难性学习。我们将此场景形式化,定义了威胁模型,并提出了六种攻击策略:标签交换(Label-Exchange)、张量交换(Tensor-Exchange)、吸引重合(Attraction-Coincident)、吸引前置(Attraction-Preceding)、排斥重合(Repulsion-Coincident)和排斥前置(Repulsion-Preceding)。吸引变体最小化中毒迭代与目标迭代标签之间的损失,在特征空间中将它们的表示拉到一起;排斥变体最大化此损失,将它们推开,使稳定性机制抵抗所需的参数偏移。在重合变体中,中毒迭代与目标迭代重合,仅使用干净的参考迭代作为标签源;在前置变体中,中毒迭代先于目标迭代,由于表示失真导致目标迭代不可学习。我们在MNIST和Split-CIFAR10数据集上,针对DER、ER-ACE和iCaRL三种CL策略,进行了超过4480次模拟评估。结果显示存在严重漏洞:攻击者可选择性阻碍可塑性,妨碍新知识的获取,同时促使先前知识的丧失,引发灾难性学习场景。
英文摘要
Continual Learning (CL) enables deep learning models to iteratively learn from a stream of data without forgetting prior knowledge. Existing adversarial research on CL primarily aims to re-enable catastrophic forgetting, attacking stability and reducing availability. We identify a novel security flaw: data manipulated by an attacker can reduce the learnability of current or upcoming iterations. We term such manipulations learning blockers, as they attack the plasticity of CL algorithms. They are particularly harmful because they are difficult to detect during training of the current iteration, since they can target iterations whose data the model has not yet encountered. When learning blockers additionally induce catastrophic forgetting, the resulting overall degradation is what we call catastrophic learning. We formalize this scenario, define a threat model and propose six attack strategies: Label-Exchange, Tensor-Exchange, Attraction-Coincident, Attraction-Preceding, Repulsion-Coincident, and Repulsion-Preceding. The Attraction variants minimize the loss between the poisoned and the victim iteration label, pulling their representations together in feature space; the Repulsion variants maximize this loss, pushing them apart so stability mechanisms resist the required parameter shift. In the Coincident variants, the poisoned and the victim iteration coincide, using a clean reference iteration only as a label source; in the Preceding variants, the poisoned iteration precedes the victim, leaving it unlearnable due to distorted representations. We evaluate on MNIST and Split-CIFAR10 against three CL strategies - DER, ER-ACE, and iCaRL - across more than 4,480 simulations. Our results demonstrate a strong vulnerability: an adversary can selectively impede plasticity to hinder the acquisition of new knowledge, while promoting loss of prior knowledge, inducing a catastrophic learning scenario.