CAM引导的显著性Cutout与基于图像的恶意软件分类
CAM-Guided Saliency Cutout and Image-Based Malware Classification
- San Jose State University(圣何塞州立大学)
- Czech Technical University in Prague(布拉格捷克技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究探究CAM引导的显著性Cutout对恶意软件图像分类的作用,通过对比不同Cutout设置在RawMal-TF和CIFAR-100数据集上的表现,发现其价值具有领域依赖性,恶意软件图像与自然图像存在差异。
AI中文摘要:
Dropout正则化通过在训练过程中移除神经网络的部分组件来减少过拟合,这一方法被广泛使用。对于卷积神经网络(CNN)而言,Cutout的作用与其有一定相似性,Cutout可作为数据增强技术实现:保留原始训练图像,同时创建移除了部分区域的额外副本。在本研究中,我们测试是否可通过高分辨率类激活映射(HiResCAM)优化Cutout的放置方式。我们对比了四种受控训练条件:无Cutout、标准随机Cutout、低显著性Cutout、高显著性Cutout。我们使用RawMal-TF数据集的灰度恶意软件图像开展实验,该数据集包含17个家族,每个家族约1000个样本;为与自然图像对比,我们还使用知名的CIFAR-100数据集进行实验。所有实验均基于ResNet18模型,训练轮次约为100轮。对于Cutout实验,我们测试了约5%、10%、20%、30%的Cutout区域,且考虑了每张原始训练图像对应M∈{4,8}个增强副本的情况。RawMal-TF数据集的实验结果显示,三种Cutout情况(随机、高显著性、低显著性)的表现均略差于无Cutout的情况;相比之下,CIFAR-100数据集在低显著性Cutout下的实验结果略有提升。这些结果表明,显著性引导Cutout的价值具有领域依赖性,恶意软件图像不应被视为与自然图像等价的对象。
英文摘要:
Dropout regularization is commonly used to reduce overfitting by removing parts of a neural network during training. For Convolutional Neural Networks (CNN), cutouts serve a somewhat analogous purpose. Cutouts can be implemented as data augmentation: the original training image is retained, and additional copies are created with regions removed. In this chapter, we test whether cutout placement can be improved by using High-Resolution Class Activation Mapping (HiResCAM). We compare four controlled training conditions: no cutout, standard random cutout, low-saliency cutout, and high-saliency cutout. We experiment using grayscale malware images from the RawMal-TF dataset (17 families with~1,000 samples per family), and for comparison to natural images, we experiment with the well-known CIFAR-100 dataset. All experiments are based on ResNet18 with~100 training epochs. For the cutout experiments, we test cutout areas of~5\%, 10\%, 20\%, and~30\%, and we consider~$M\in\{4,8}$ augmented copies per original training image. The RawMal-TF results are slightly worse for all three cutout cases (random, high and low saliency) as compared to no cutouts. In contrast, our CIFAR-100 experimental results improve slightly under low-saliency cutout. These results suggest that the value of saliency-guided cutout is domain dependent, and that malware images should not be treated as equivalent to natural images.