基于图像技术与集成软投票的恶意软件分类
Image-Based Techniques and Ensemble Soft Voting for Malware Classification
- Department of Computer Science, San Jose State University(圣何塞州立大学计算机科学系)
- Faculty of Information Technology, Czech Technical University in Prague(布拉格捷克理工大学信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究利用集成软投票框架,结合手工特征、预训练嵌入和自定义CNN特征,对八种图像转换的恶意软件进行分类,最佳集成达到80.2%准确率,较最优单模型提升2.4个百分点。
AI中文摘要:
在本章中,我们研究使用集成学习框架和软投票策略进行基于图像的恶意软件家族分类。我们考虑了使用八种不同转换策略将恶意软件二进制文件转换为图像的情况。对这些图像应用了三条互补的特征提取路径:手工设计的描述符,结合了方向梯度直方图(HOG)和Haralick纹理特征以及38个统计特征;从三个预训练神经网络(VGG16、ResNet50和ViT-B/16)获得的密集嵌入,其中每个预训练模型用作冻结的特征提取器,并移除其分类头;以及从直接在恶意软件图像上训练的自定义卷积神经网络(CNN)导出的512维嵌入。三种特征提取技术中的每一种都使用机器学习分类器在所有八种图像转换类型上进行评估。最佳个体结果是手工特征77.8%的准确率,预训练神经网络路径73.8%,自定义CNN路径74.8%。然后我们考虑了各种软投票集成策略,发现表现最佳的软投票池——由在专用验证集上选择的十五个投票者组成——在考虑的17个恶意软件家族中实现了80.2%的准确率,比最佳个体模型统计显著提高了2.4个百分点。定量多样性分析确认了不同特征表示是互补的,其中手工描述符是最强的贡献者。
英文摘要:
In this chapter, we investigate image-based malware family classification using an ensemble learning framework and a soft voting strategy. We consider malware binaries that have been converted into images using eight distinct conversion strategies. Three complementary feature extraction tracks are applied to these images: handcrafted descriptors combining Histogram of Oriented Gradients (HOG) and Haralick texture features along with 38 statistical features; dense embeddings obtained from three pretrained neural networks (VGG16, ResNet50, and ViT-B/16), where each pretrained model is used as a frozen feature extractor with its classification head removed; and 512-dimensional embeddings derived from a custom Convolutional Neural Network (CNN) trained directly on the malware images. Each of the three feature extraction techniques is evaluated with machine learning classifiers across all eight image conversion types. The best individual results are 77.8% accuracy for the handcrafted features, 73.8% for the pretrained neural network track, and 74.8% for the custom CNN track. Then we consider various soft voting ensemble strategies, and we find that the best-performing soft voting pool--consisting of fifteen voters selected on a dedicated validation split--achieves 80.2% accuracy across the 17 malware families under consideration, a statistically significant improvement of 2.4 percentage points over the best individual model. A quantitative diversity analysis confirms that the different feature representations are complementary, with the handcrafted descriptors being the strongest contributors.