arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34456cs.CR

打破Windows恶意软件检测:问题空间对抗鲁棒性的全面评估

Breaking Windows Malware Detection: A Comprehensive Evaluation of Problem-Space Adversarial Robustness

Mashal Zainab, Salijona Dyrmishi, Hamid Bostani, Lorenzo Cavallaro, Maxime Cordy

首次发表
浏览论文内容

中文总结 AI 辅助

本研究统一大规模评估了九种逃逸攻击对八款Windows恶意软件检测器的对抗鲁棒性,发现脆弱性取决于模型表示和攻击类型,最强攻击用更少变换获更高成功率,且对抗加固难以跨攻击迁移。

中文摘要 AI 辅助

问题空间逃逸攻击已暴露出基于机器学习的恶意软件检测器的关键弱点;然而,对其评估在模型、数据集和攻击方法上仍然零散,且常常忽视可执行性和功能保持等特定领域要求。我们通过一项统一的大规模评估来弥补这一空白,在保持可执行性的条件下,对八款Windows恶意软件检测器(包括七个开源模型和一个商业检测器)评估了九种最先进的逃逸攻击。我们的研究分析了攻击有效性、互补性、可迁移性和对抗性加固,以从互补维度评估鲁棒性。我们表明,检测器的脆弱性在很大程度上取决于模型表示和攻击类型:原始字节检测器特别容易受到几类问题空间操纵的影响,但没有一个检测器家族在所有攻击下都表现出统一的鲁棒性。重要的是,有效性并不能仅由变换空间大小来解释:最强的攻击可以使用更少的独特变换并集中于一小部分高影响操纵,从而获得显著更高的成功率。我们进一步表明,两种互补攻击足以覆盖其余评估攻击所产生的约99%的对抗样本。可迁移性表现出与直接攻击成功不同的模式:直接成功率低的攻击可以产生高度可迁移的逃逸。最后,对抗性加固高度依赖于攻击和模型:鲁棒性增益往往无法跨攻击迁移,甚至可能增加对未见攻击的敏感性。这些发现凸显了当前恶意软件鲁棒性评估的局限性,建立了全面的经验基线,并阐明了有效性、可迁移性和防御鲁棒性之间的关系。

英文摘要

Problem-space evasion attacks have exposed critical weaknesses in machine learning-based malware detectors; yet, their evaluation remains fragmented across models, datasets, and attack methodologies, often neglecting domain-specific requirements such as executability and functionality preservation. We address this gap with a unified, large-scale evaluation of nine state-of-the-art evasion attacks against eight Windows malware detectors, including seven open-source models and one commercial detector, under executability-preserving conditions. Our study analyzes attack effectiveness, complementarity, transferability, and adversarial hardening to evaluate robustness along complementary dimensions. We show that detector vulnerability depends strongly on both model representation and attack type: raw-byte detectors are particularly susceptible to several classes of problem-space manipulation, but no detector family is uniformly robust across all attacks. Importantly, effectiveness is not explained by transformation-space size alone: the strongest attacks can achieve substantially higher success while using fewer distinct transformations and concentrating on a small set of high-impact manipulations. We further show that two complementary attacks are sufficient to cover approximately 99% of the adversarial examples produced by the remaining evaluated attacks. Transferability exhibits a different pattern from direct attack success: attacks with low direct success can produce highly transferable evasions. Finally, adversarial hardening is highly attack- and model-dependent: robustness gains often fail to transfer across attacks and can even increase susceptibility to unseen attacks. These findings highlight limitations in current malware robustness evaluations, establish a comprehensive empirical baseline, and clarify relationships between effectiveness, transferability, and defense robustness.

发表机构

  • University of Luxembourg(卢森堡大学)
  • University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑