arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强化学习启发的黑盒对抗攻击用于计算机视觉

Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision

Florian Krone, Elena Hoemann, Sven Hallerbach

arXiv 2609.24249首次发表:更新:

发表机构

German Aerospace Center (DLR)(德国航空航天中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对黑盒场景下神经网络易受扰动攻击的问题,提出基于强化学习的RIBA方法,通过少量查询生成对抗样本,在Cifar10和ImageNet上显著减少查询次数,并匹配白盒攻击性能。

AI 中文摘要

神经网络,无论是基于卷积还是基于Transformer,对于现代计算机视觉系统至关重要。然而,它们容易受到微小扰动的影响,这些扰动对人类几乎不可察觉,却会显著改变模型的预测。这些对抗攻击通常被认为是对神经网络在安全关键应用中部署的重大威胁。大多数攻击采用白盒威胁模型,因此需要完全访问目标模型,这使得它们在实践中不切实际。我们提出了一种在更现实的黑盒威胁模型下的新颖方法,该方法利用强化学习的概念来优化具有不可微分目标模型的扰动。强化学习算法已经被优化为查询高效,使其成为设计黑盒对抗攻击的理想起点。我们通过将我们的强化学习启发的黑盒对抗攻击(RIBA)与Cifar10和ImageNet数据集上不同模型的最先进攻击进行比较,展示了RIBA仅使用少量对目标模型的查询就能生成对抗扰动的成功。RIBA在Cifar10上针对ResNet-18生成对抗图像所需的中位查询次数减少了25.4%,在ImageNet上欺骗Vit-B/16模型所需的中位查询次数减少了22.5%。此外,我们证明RIBA在对抗训练模型上可以达到白盒攻击的性能。

英文摘要

Neural networks, both convolution or transformer based, are essential for modern computer vision systems. However, they are vulnerable to small perturbations, almost imperceptible to humans, which significantly alter the model's prediction. These adversarial attacks are often considered to be a significant threat to the implementation of neural networks in safety-critical applications. Most attacks utilize the white-box threat model and therefore require full access to the target model, making them unrealistic to use in practice. We propose a novel approach under the more realistic black-box threat model that utilizes concepts from reinforcement learning to optimize perturbations with a non-differentiable target model. Reinforcement learning algorithms have already been optimized to be query efficient, making them an ideal starting point when designing black-box adversarial attacks. We show the success of our reinforcement learning inspired black-box adversarial attack (RIBA) in generating adversarial perturbations using only a small number of queries to the target model, by comparing it to state of the art attacks on different models on the Cifar10 and ImageNet data sets. RIBA takes $25.4\%$ fewer median queries to generate attacked images against a ResNet-18 on Cifar10 and $22.5\%$ fewer median queries to fool a Vit-B/16 model on ImageNet. Additionally, we demonstrate that RIBA can match the performance of white-box attacks on an adversarially trained model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑