REPLICANT:学习用于规避和加固恶意软件检测器的策略
REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
- University College London(伦敦大学学院)
- KU Leuven(鲁汶大学)
- The Alan Turing Institute(阿兰·图灵研究所)
- King’s College London(伦敦国王学院)
- Core64
- Devotion AI Labs(Devotion AI实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Replicant是深度强化学习框架,在仅标签黑盒威胁模型下学习恶意软件规避策略,在7种Android恶意软件检测器上攻击成功率达78.8%,还可用于提升检测器鲁棒性。
AI中文摘要:
为确定基于机器学习的恶意软件检测在现实世界中的有效性,评估其对抗高能力攻击者的鲁棒性至关重要。然而,最先进的攻击方法并未有效建模现实攻击者,因为它们通常假设能获取特权信息,如目标的训练数据、特征空间或置信分数。在本研究中,我们提出Replicant,这是一个深度强化学习框架,可在严格的仅标签黑盒威胁模型下学习规避的现实任务。Replicant学习可复用的策略,包括如何修改恶意软件样本以及何时查询目标,该策略可跨样本、检测器和特征空间迁移。在7种Android恶意软件检测器和3种特征空间上,Replicant是最强且查询效率最高的方法,实现了78.8%的平均攻击成功率,相比最先进方法的相对提升为20.9%-39.2%。此外,当用于对抗训练时,Replicant还能生成具有更具泛化性鲁棒性的检测器,其性能优于最先进方法。通过Replicant,我们证明学习规避任务不仅能带来更强的攻击性能,更关键的是,能为加固恶意软件检测器提供更好的信号。
英文摘要:
To determine the real-world effectiveness of machine learning based malware detection, it is vital to evaluate its robustness against highly capable adversaries. However, state-of-the-art attacks do not effectively model realistic adversaries, as they often assume access to privileged information such as the training data, feature space, or confidence scores of the target. In this work, we present Replicant, a deep reinforcement learning framework that learns the realistic task of evasion under a strict label-only black-box threat model. Replicant learns a reusable policy on how to modify a malware sample and when to query the target, which transfers across samples, detectors, and feature spaces. Across seven Android malware detectors and three feature spaces, Replicant is the strongest and most query-efficient approach achieving a mean attack success rate of 78.8%, a relative improvement of 20.9%-39.2% over the state-of-the-art. Furthermore, when used for adversarial training, Replicant also outperforms the state-of-the art by producing detectors with more generalizable robustness. With Replicant we demonstrate that learning the task of evasion not only results in stronger attack performance but, crucially, provides a better signal for hardening malware detectors.