arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Malformer:一种基于Transformer的多模态恶意软件检测器

Malformer: A Multi-Modal Malware Detector Using Transformers

Samuel Howard, Kshitiz Aryal, Mahmoud Abdelsalam, Maanak Gupta, Andrew Wheeler, Pradip Kunwar

arXiv 2608.19052首次发表:更新:

AI 中文总结

本研究提出四模态恶意软件检测模型Malformer,融合文本、图像、图形、音频表示,采用多模态Transformer融合,在201549个样本数据集上准确率98.3%,优于单/双模态检测器,为恶意软件防御提供稳健基础。

AI 中文摘要

传统恶意软件检测系统依赖单一恶意软件表示,往往无法识别新型威胁。这些恶意软件二进制文件的表示(也称为模态)无法为模型提供足够信息以区分所有样本,且单一表示会引入新的失效模式,部分模态提取依赖反汇编的成功与否。过往研究要么整合额外模态,要么采用更具区分性的表示进行分类。本研究提出Malformer,一种四模态恶意软件检测模型,整合Windows可执行文件的文本、图像、图形和音频表示。我们证明多模态Transformer融合可提升Windows恶意软件检测器的性能,优于单模态和双模态检测器。Malformer采用两个RoBERTa编码器、经修改的Vision Transformer(用于图像数据)、WavLM(用于音频数据),并结合自适应损失加权方案融合各模态特定表示。在包含201549个二进制样本的数据集上评估,Malformer准确率达98.3%,F1分数为0.9833,较单模态基线和双模态检测器提升4.6至17.6个百分点。Malformer表明多模态融合为应对日益增长的恶意软件威胁规模提供了有前景的基础,为防御者提供通用且稳健的检测能力。

英文摘要

Traditional malware detection systems that rely on a single representation of malware often fail to identify novel threats. These representations of malware binaries, also known as modalities, do not provide the models with sufficient information to discriminate among all samples. Additionally, individual representations introduce new failure modes, with some modality extraction being dependent upon the success of disassembling. Past works have integrated either additional modalities or more discriminative representations for classification. In this work, we present Malformer, a quadrimodal malware detection model that incorporates text, image, graph, and audio representations of Windows executables. We demonstrate that multimodal transformer fusion can enhance the performance of Windows malware detectors over that of unimodal and bimodal detectors. Malformer employs a combination of two RoBERTa encoders paired with a modified Vision Transformer for image data, WavLM for audio data, and an adaptive loss-weighting scheme to fuse modality-specific representations. Evaluated on a dataset of 201,549 binary samples, Malformer achieved 98.3% accuracy and an F1 score of 0.9833, outperforming both unimodal baselines and bimodal detectors by 4.6-17.6 percentage points. Malformer demonstrates that multimodal fusion provides a promising foundation for countering the growing scale of malware threats, equipping defenders with generalized and resilient detection capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑