发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种通过向神经网络提供TTP信息并训练其识别具体恶意TTP的方法,实现精确恶意软件检测,优于不利用TTP信息的模型,尤其在罕见TTP和对抗攻击场景下表现突出。
AI 中文摘要
机器学习方法,尤其是神经网络,现在常被用于网络流量中的恶意软件检测。尽管这些方法非常有效,但基于此类方法的系统通常(i)纯粹是数据驱动的,忽略了关于可能使用的战术、技术和程序(TTP)的大量现有知识,因此(ii)不够精确,因为它们要么无法将恶意活动与 TTP 使用相关联,要么即使能够关联,也无法解释哪个 TTP 被恶意使用。在本文中,我们证明了通过(i)向神经网络模型提供关于任何给定样本所使用的 TTP 的信息,以及(ii)教导神经网络不仅检测整体恶意活动,而且检测哪些特定 TTP 被恶意使用,可以精确地检测恶意软件。我们表明,我们的方法始终优于三种替代模型,这些模型要么不利用 TTP 信息,要么未被教导检测 TTP 的恶意使用,或两者兼有。此外,我们表明我们的方法(i)在检测利用罕见使用的 TTP 的恶意软件方面特别有益,这种情况对其他系统尤其具有挑战性;(ii)允许 TTP 对 TTP 的调优,进一步提高其检测 TTP 恶意使用的能力;(iii)在广泛场景中始终优于其他系统,包括依赖有限训练数据或遭受对抗性攻击时。
英文摘要
Machine learning methods, and especially neural networks, are now routinely used for malware detection in network traffic. Though very effective, systems based on such methods often (i) are purely data-driven, ignoring the substantial body of available knowledge about the tactics, techniques, and procedures (TTPs) possibly used, and, consequently (ii) are not precise, since they either cannot correlate malicious activity with TTP usage, or if they do, they are unable to explain which TTP has been maliciously used. In this paper we demonstrate that it is possible to precisely detect malware by (i) providing the neural network model with information about the TTPs used by any given sample, and (ii) teaching the neural network to detect not just the malicious activity as a whole, but which specific TTPs are maliciously used. We show that our approach consistently outperforms the three alternative models, which either do not exploit TTP information, or which are not taught to detect the malicious usage of TTPs, or both. Moreover, we show that our approach (i) is particularly beneficial in detecting malware that utilises rarely-used TTPs, a scenario which is particularly challenging for the other systems; (ii) allows for TTP by TTP tuning, further improving its ability to detect the malicious usage of TTPs; (iii) consistently outperforms other systems across a wide-range of scenarios, including when relying on limited training data or when subjected to adversarial attack.