DDQN-MLP:一种可解释且对抗鲁棒的DRL引导自适应学习框架用于勒索软件检测
DDQN-MLP: An Explainable and Adversarially Robust DRL-Guided Adaptive Learning Framework for Ransomware Detection
浏览论文内容
中文总结 AI 辅助
提出DDQN-MLP框架,利用双深度Q网络自适应加权指导MLP进行勒索软件检测,在2000样本数据集上达到99.30%准确率,兼具可解释性与对抗鲁棒性。
中文摘要 AI 辅助
勒索软件检测仍然具有挑战性,因为现代变体表现出多样化、规避性以及部分类似良性的行为,这些行为破坏了固定的监督学习目标。本研究提出了DDQN-MLP,一种训练时的深度强化学习框架,利用Windows 11沙箱遥测数据进行行为勒索软件检测。双深度Q网络(DDQN)作为离散的自适应样本权重控制器,通过观察批次级别的损失和预测置信度动态,分配样本重要性权重以指导轻量级多层感知器(MLP)。训练后,DDQN被丢弃,仅保留高效的MLP用于部署。该框架在一个包含2,000个可执行文件配置文件的平衡数据集上进行了5折分层交叉验证评估,其中包括来自30个家族的1,000个勒索软件样本和1,000个良性样本。DDQN-MLP达到了99.30%的准确率、0.9930的F1分数和0.9991的ROC-AUC,优于传统的静态加权、焦点损失和替代DRL变体。使用SHAP和LIME评估了可解释性,并使用了SHAP梯度对齐诊断来评估特征归因与模型敏感性之间的一致性。跨多个扰动水平的白盒对抗测试进一步表明,对抗训练提高了特征空间鲁棒性,而不会降低干净数据的准确性。结果表明,DDQN-MLP为高通量勒索软件检测提供了一种准确、可解释、鲁棒且计算高效的框架。
英文摘要
Ransomware detection remains challenging because modern variants exhibit diverse, evasive, and partly benign-like behaviors that undermine fixed supervised learning objectives. This study proposes DDQN-MLP, a training-time deep reinforcement learning framework for behavioral ransomware detection using Windows 11 sandbox telemetry. A Double Deep Q-Network (DDQN) acts as a discrete adaptive sample-weighting controller by observing batch-level loss and prediction-confidence dynamics and assigning sample-importance weights to guide a lightweight Multilayer Perceptron (MLP). After training, the DDQN is discarded, leaving only the efficient MLP for deployment. The framework was evaluated using 5-fold stratified cross-validation on a balanced dataset of 2,000 executable profiles comprising 1,000 ransomware samples from 30 families and 1,000 benign samples. DDQN-MLP achieved 99.30% accuracy, an F1-score of 0.9930, and an ROC-AUC of 0.9991, outperforming conventional static weighting, focal-loss, and alternative DRL variants. Explainability was assessed using SHAP and LIME, together with a SHAP-gradient alignment diagnostic for evaluating consistency between feature attribution and model sensitivity. White-box adversarial testing across multiple perturbation levels further showed that adversarial training improved feature-space robustness without reducing clean-data accuracy. The results demonstrate that DDQN-MLP provides an accurate, explainable, robust, and computationally efficient framework for high-throughput ransomware detection.
发表机构
- Charles Sturt University(查尔斯史都华大学)
- University of New South Wales(新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。