arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种混合恶意软件检测方法:将少样本模型无关元学习与自编码器相结合

A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders

Emmanuela Andam, Yasir Abbas Zaidi, Abdelali Hadir, Emmanuel Grant, Naima Kaabouch

arXiv 2610.01949首次发表:更新:

发表机构

University of North Dakota; Hassan II University(北达科他大学; 哈桑二世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种结合自编码器特征提取器与模型无关元学习分类器的混合深度学习框架,用于少样本勒索软件检测,在Ransomware Dataset 2024上实现高准确率、F1分数和MCC值,验证了其在数据稀缺场景下的鲁棒性和有效性。

AI 中文摘要

勒索软件已成为一个主要的网络安全威胁,其攻击事件在关键行业的频率和影响均不断增加。这些攻击通常通过钓鱼邮件、恶意下载或利用软件漏洞获取系统访问权限来发起。一旦进入系统,恶意软件会加密文件并要求支付赎金(通常以加密货币形式)以换取解密密钥。传统的检测方法往往难以应对新颖或稀少的样本,使系统易受攻击。为解决这些问题,本文提出了一种混合深度学习框架,该框架将自编码器特征提取器(AFE)与模型无关元学习(MAML)分类器相结合,用于少样本恶意软件检测。AFE生成紧凑的潜在特征,以减少噪声和维度,而MAML分类器则利用有限的标记数据快速适应新威胁。在Ransomware Dataset 2024上进行的实验证明了该框架在二分类任务中的有效性。在一到五十样本(one to fifty shot)的设置下,所提出的模型始终达到高准确率、F1分数和马修斯相关系数(MCC)值,即使在极端稀缺情况下也能保持可靠的分类性能。这些结果突显了该模型在有限数据场景下的鲁棒性和有效性,展示了将特征提取与元学习相结合以增强对恶意软件抵御能力的潜力,特别是在医疗保健、制造业和公共基础设施等行业,这些领域的网络攻击可能造成重大的运营和财务中断。

英文摘要

Ransomware has emerged as a major cybersecurity threat, with incidents increasing in frequency and impact across critical sectors. These attacks are typically launched through phishing emails, malicious downloads, or exploitation of software vulnerabilities to gain system access. Once inside, the malware encrypts files and demands a ransom, often in cryptocurrency, for the decryption key. Conventional detection methods often struggle with novel or scarce samples, leaving systems vulnerable. To address these challenges, this paper proposes a hybrid deep learning framework that combines an Autoencoder Feature Extractor (AFE) with a Model Agnostic Meta Learning (MAML) classifier for few shot malware detection. The AFE generates compact latent features that reduce noise and dimensionality, while the MAML classifier rapidly adapts to new threats using limited labeled data. Experiments conducted on the Ransomware Dataset 2024 demonstrate the effectiveness of the framework in binary classification tasks. Across one to fifty shot settings, the proposed model consistently achieves high accuracy, F1 score, and Matthews Correlation Coefficient values, maintaining reliable classification even under extreme scarcity. These results highlight the model's robustness and effectiveness in adapting to limited data scenarios, demonstrating the potential of combining feature extraction with meta learning to enhance resilience against malware, particularly in sectors such as healthcare, manufacturing, and public infrastructure, where cyberattacks can cause significant operational and financial disruption.

CommentsAccepted at 2025 Cyber Awareness and Research Symposium (CARS). This is the author's accepted manuscript

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑