arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规避攻击:对抗性噪声如何绕过机器学习分类器

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

Parker Hummel, Ryne Skabo, Muhammad Abusaqer

arXiv 2610.00136首次发表:更新:

发表机构

Minot State University(迈诺特州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过MNIST图像分类和SMS文本分类实验,展示对抗性噪声对机器学习分类器的规避攻击效果,发现图像分类脆弱性显著而文本分类相对稳健,强调鲁棒性需实证检验。

AI 中文摘要

本文对图像分类和文本分类中的规避攻击进行了可复现的、教学性的研究。在MNIST数据集上训练的紧凑卷积网络达到了98.63%的干净测试准确率,并在两种白盒攻击下进行了评估。在FGSM攻击下,当ε=0.15时准确率降至60.20%,当ε=0.30时降至1.72%;在PGD攻击下,准确率分别降至32.47%和0.41%,而位深度缩减防御仅恢复了部分损失。在第二个实验中,在SMS垃圾短信集上微调的DistilBERT达到了98.75%的准确率和94.96%的F1分数,但一系列受控的预定义扰动(字符替换、空白噪声和良性后缀)在大多数展示的示例中仅产生了适度的概率变化,且没有发生从垃圾邮件到正常邮件的翻转。对抗性脆弱性强烈依赖于模态:MNIST实验是明确的规避演示,而文本实验是受控的鲁棒性评估。鲁棒性必须通过实证测试,而不能仅从干净准确率推断。

英文摘要

This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at $ε$ = 0.15 and 1.72% at $ε$ = 0.30; under PGD it fell to 32.47% and 0.41%, and a bit-depth-reduction defense recovered only part of the loss. In the second experiment, DistilBERT fine-tuned on the SMS Spam Collection reached 98.75% accuracy and a 94.96% F1-score, but a controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed examples and no flip from spam to ham. Adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation. Robustness must be tested empirically rather than inferred from clean accuracy.

CommentsPresented at the 58th Midwest Instruction and Computing Symposium (MICS 2026), Eau Claire, WI, March 27 to 28, 2026. 14 pages, 7 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑