arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种新型对抗样本:衡量人类与模型的差距及其与分布外检测的关系

A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection

Ali Borji

arXiv 2607.22722首次发表:更新:

发表机构

BrainChip Inc.(脑芯片公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究一种新型对抗样本,通过在MNIST、CIFAR-10和ImageNet上实验,回答人类与模型差距、OOD检测工具及现有防御能否应对该样本三个问题,发现人类表现比模型差,现有检测工具失效,经典防御无效,且攻击对低层次纹理破坏更快。

AI 中文摘要

几乎所有对抗攻击都添加难以察觉的扰动来愚弄模型。本文研究相反情况:添加大的、清晰可见的扰动,使模型保持正确预测,而人类无法识别图像。先前工作表明可大规模生成此类样本,但有三个问题未测试:人类表现是否真比模型差、标准分布外(OOD)检测和校准工具能否检测到、现有防御能否减轻影响。本文在MNIST、CIFAR-10和ImageNet上回答了这三个问题。独立识别代理在CIFAR-10上降至约49%,而模型保持100%,人类也证实了差距;基于置信度和能量的OOD检测器和校准无法检测到,特征空间马氏距离检测器能检测但会被自适应攻击者规避;没有经典防御能降低攻击成功率。机理分析表明攻击破坏低层次纹理比边缘/形状结构快得多。

英文摘要

Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible perturbation that causes the model to keep its original, correct prediction, even though a human would no longer recognize the image. Prior work showed such examples can be generated at scale but left three questions untested: whether humans really perform worse than the model, whether standard out-of-distribution (OOD) detection and calibration tools catch it, and whether existing defenses mitigate it. We answer all three on MNIST, CIFAR-10, and ImageNet. (i) An independent recognizer proxy drops to ~49% on CIFAR-10 while the model stays at 100% -- a gap a small human pilot (N=5) corroborates directly and that is not explained by signal loss (a matched-magnitude Gaussian control degrades recognizability faster); a CLIP zero-shot proxy confirms the gap at ImageNet scale too. (ii) Confidence- and energy-based OOD detectors and calibration are structurally blind (0% detection, ECE ~= 0), while a feature-space Mahalanobis detector flags 100% -- but is evaded by an adaptive attacker at no cost to success. (iii) No classical defense, including adversarial training (45% robust accuracy), reduces attack success (correlation with large-epsilon_l resistance r ~= 0). A mechanistic analysis further shows the attack destroys low-level texture far faster than edge/shape structure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑