arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12077cs.CR

基于Grad-CAM与混合学习模型的恶意软件图像变换方法对比

A Comparison of Malware Image Transformations Using Grad-CAM and Hybrid Learning Models

Vibha Bhavikatti, Mark Stamp

首次发表
浏览论文内容

中文总结 AI 辅助

本研究对比8种恶意软件图像变换方法,采用Grad-CAM分析,发现准确率与解释保真度不一致,基于Grad-CAM特征的随机森林模型在17个恶意软件家族上准确率达0.777,优于此前基准。

中文摘要 AI 辅助

近期研究表明,二进制转图像的表示方式可实现基于机器学习的有效恶意软件检测与分类,但性能会因二进制转图像所采用的技术存在显著差异。此外,图像类模型的可解释性在恶意软件领域尚未得到充分探索。本研究采用梯度加权类激活映射(Grad-CAM)作为可解释人工智能(XAI)工具,用于分析从恶意软件样本衍生出的8种不同图像类型;提供了Grad-CAM热图的定量保真度与稳定性指标,并将这些热图与高分辨率类激活映射(HiResCAM)进行对比。研究还表明,Grad-CAM热图可为恶意软件分类提供有用信息,具体而言,通过MobileNetV2卷积神经网络(CNN)模型从Grad-CAM图像提取特征后训练的随机森林模型,在17个恶意软件家族上的测试准确率达0.777,超过了该数据集此前0.750的基准值。本研究的关键发现是,对于所考虑的恶意软件图像变换方法,准确率与解释保真度并不一致,例如,生成最具保真度解释的图像变换技术仅能产生中等水平的准确率。

英文摘要

Recent studies have shown that binary-to-image representations can enable effective machine learning-based results for malware detection and classification. However, performance can vary significantly, depending on the technique used to convert binaries to images. Furthermore, the explainability and interpretability of image-based models is largely unexplored within the malware domain. In this research, we employ Gradient-weighted Class Activation Maps (Grad-CAM) as an eXplainable AI (XAI) tool, which we use to analyze eight distinct image types derived from malware samples. We provide quantitative faithfulness and stability metrics for Grad-CAM heatmaps and we compare these heatmaps to High-Resolution Class Activation Mappings (HiResCAM). We also show that Grad-CAM heatmaps can provide useful information for malware classification. Specifically, we show that a Random Forest model trained on features extracted from Grad-CAM images via a MobileNetV2 Convolutional Neural Network (CNN) model achieves a test accuracy of 0.777 across 17 malware families, exceeding a previous benchmark of 0.750 for this same dataset. A key finding of this research is that for the malware image transformations considered, accuracy and explanation faithfulness do not coincide, e.g., image transformation techniques that produce the most faithful explanations yield only mid-tier accuracy.

补充信息

↑