发表机构
Neurint, LLC; U.S. Army Cyber Corps; Northwestern State University of Louisiana; Automata; AGH University of Krakow(Neurint有限责任公司; 美国陆军网络部队; 路易斯安那西北州立大学; Automata公司; 克拉科夫AGH大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉数据中自然出现的类似后门的统计信号(统计对手),通过分析ImageNet找相关模式,用统计控制去除随机相关性,证明其可改变模型预测,表明漏洞受数据集结构驱动,建议审核数据集时重视虚假结构。
AI 中文摘要
特定模型的对抗攻击已被广泛研究。我们研究了一种不同的失败模式:视觉数据中自然出现的统计信号,其表现类似后门触发因素但并非恶意插入。我们将这些信号称为统计对手。我们分析ImageNet以找到与某些标签紧密相关的模式,然后使用统计控制去除候选信号中的随机相关性。最后,我们证明这些信号直接且可预测地改变模型预测。这些统计对手比一般损坏更具针对性且能跨不同模型架构转移。这表明某些漏洞由数据集结构和分布驱动而非单个模型特性。我们得出结论,普通数据集即使没有中毒也可能包含可利用的对抗表面,并建议数据集审核应将虚假结构不仅视为偏差或可解释性失败的来源,还应视为视觉模型潜在的攻击表面。
英文摘要
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signals in vision data that can behave as backdoor-like triggers without being maliciously inserted. We call these signals statistical adversaries. We analyse ImageNet to find patterns that are strongly linked to certain labels. We then use statistical controls to remove random correlations from our candidate signals. Finally, we demonstrate that these signals directly and predictably alter model predictions. These statistical adversaries are more targeted than generic corruptions and transfer across different model architectures. This suggests that some vulnerabilities are driven by dataset structure and distribution rather than a single model's idiosyncrasies. We conclude that ordinary datasets can contain exploitable adversarial surfaces even in the absence of poisoning, and suggest that dataset audits should treat spurious structure not only as a source of bias or interpretability failure, but also as a latent attack surface for vision models.