arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

增强资源受限环境中的多类恶意软件分类

Enhancing Multiclass Malware Classification in Resource-Constrained Environments

Abdul Khalek Alve, Alif Rahman, Saadman Zaman, Sazzad Hossen Himel, Muhammad Iqbal Hossain

arXiv 2609.27950首次发表:更新:

发表机构

BRAC University(BRAC大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对资源受限环境中多类恶意软件分类准确率低且计算开销大的问题,提出结合LightGBM与随机森林的轻量级模型,通过SMOTE和SOM-US平衡数据、遗传算法选择特征,在家族和个体分类上分别达到91.2%和78.7%的准确率,优于现有方法。

AI 中文摘要

勒索软件、间谍软件、木马等多类恶意软件攻击的出现,对网络安全构成了日益严重且严峻的威胁,尤其是在物联网设备等资源受限的环境中。现有的机器学习模型在二分类恶意软件检测中已实现近乎完美的准确率,但在恶意软件家族和个体恶意软件的分类方面表现不足。此外,这些多类恶意软件攻击的复杂性给资源受限环境中的检测带来了重大挑战,因为多类检测通常需要较高的计算能力。本研究通过提高多类恶意软件分类的检测准确性,并开发一种能够在资源受限设备上高效运行的轻量级模型,弥合了这一差距。在本文中,我们提出了一种鲁棒、轻量级的机器学习模型,采用LightGBM分类器,结合SMOTE过采样和SOM-US欠采样技术进行数据平衡,并通过遗传算法进行精心设计的特征选择。该模型在相同数据集上的表现优于当前最先进的模型,在恶意软件家族分类(4类)和个体恶意软件类型分类(16类)中分别达到了89.1%和76%的准确率。从而在资源受限环境中保持了分类准确性和计算效率之间的平衡。此外,我们提出了另一个使用随机森林分类器的模型,在恶意软件家族分类中准确率为91.2%,在个体恶意软件分类中准确率为78.7%。与当前最先进的模型相比,显著提升了准确性。

英文摘要

The emergence of multi-class malware attacks such as ransomware, spyware, trojans, etc., presents an increasing and serious threat to cybersecurity, particularly in resourceconstrained environments like IoT devices. Existing machine learning models have achieved nearly perfect accuracy in binary malware classification but fall short in terms of classifying malware families and individual malware. Additionally, the complexity of these multi-class malware attacks presents a significant challenge of detection in resource-constrained environments, as multi-class detection usually requires high computational capability. This research bridges the gap by enhancing the detection accuracy of multi-class malware classification as well as developing a lightweight model that can run efficiently on resource-constrained devices. In this paper, we propose a robust, lightweight machine learning model featuring LightGBM classifier with SMOTE oversampling and SOM-US undersampling techniques for data balancing, as well as well-engineered feature selection through Genetic Algorithm. The model performed better than the current state-of-the-art models developed on the same dataset in both malware family classification (4 classes) and individual malware type classification (16 classes) with accuracy of 89.1% and 76% respectively. Thus, maintaining a balance between classification accuracy and computational efficiency in resource-constrained environments. Furthermore, we propose another model using Random Forest classifier with an accuracy of 91.2% in malware family classification and 78.7% in individual malware classification. Demonstrating a significant enhancement in terms of accuracy from the current state-of-the-art models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑