arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

恶意软件分类模型的概念漂移检测与自适应重训练

Concept Drift Detection and Adaptive Retraining of Malware Classification Models

Christofer Washington Berruz Chungata, Martin Jurecek, Katerina Potika, William B. Andreopoulos, Mark Stamp

arXiv 2608.13465首次发表:更新:

发表机构

San Jose State University; Czech Technical University in Prague(圣何塞州立大学; 布拉格捷克技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对恶意软件分类模型,提出基于OCSVM等的概念漂移检测方法,结合漂移感知重训练策略,可在保持准确率的同时提升训练效率,且OCSVM方法表现更优。

AI 中文摘要

概念漂移指数据的统计特性随时间发生变化,与训练学习模型所用数据存在差异。恶意软件检测或分类的机器学习模型极易因概念漂移导致性能下降,因为攻击者会不断修改现有恶意软件。本章分析两种基于机器学习的自动概念漂移检测方法:一种是基于单类支持向量机(OCSVM)的新颖方法,另一种是基于小批量K均值(MK-Means)的已研究技术;同时还考虑了最大均值差异(MMD)这一用于检测多维数据变化的统计技术作为对比。我们开展大量实验,对比四种学习模型的有效性:多层感知机、随机森林、支持向量机及极端梯度提升。对每种模型,我们考虑三种不同场景:无模型重训练的静态场景、无论是否存在概念漂移均持续重训练的周期性场景、仅在检测到概念漂移时才重训练的漂移感知场景。在漂移感知场景下,我们采用帕累托前沿分析,分析准确率与训练效率间的权衡关系。研究发现,三种概念漂移检测技术的分类准确率可与周期性重训练相当,同时在需重训练的模型数量方面效率显著更高;此外,基于OCSVM技术的漂移感知重训练通常优于MK-Means和MMD方法。总体而言,这些结果为我们能够准确检测恶意软件分类模型中的概念漂移提供了有力证据。

英文摘要

Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning model. Machine learning models for malware detection or classification are particularly susceptible to performance degradation caused by concept drift, as attackers constantly modify existing malware. In this chapter, we analyze two machine learning-based approaches to automated concept drift detection-a novel approach based on One-Class Support Vector Machines (OCSVM) and a previously-studied technique based on Minibatch K-Means (MK-Means). For comparison we also consider Maximum Mean Discrepancy (MMD), a statistical technique for detecting changes in multidimensional data. We conduct an extensive series of experiments comparing the effectiveness of four learning models, namely, Multilayer Perceptron, Random Forest, Support Vector Machines, and eXtreme Gradient Boosting. For each of these models, we consider three distinct scenarios: A static scenario where no model retraining occurs, a periodic scenario where models are constantly retrained irrespective of concept drift, and a drift-aware scenario where models are only retrained when concept drift is detected. Under the drift-aware scenario, we analyze the tradeoff between accuracy and training efficiency using Pareto Front analysis. We find that all three concept drift detection techniques achieve classification accuracy comparable to periodic retraining, while offering substantially greater efficiency in terms of the number of models that must be retrained. In addition, drift-aware retraining based on our OCSVM technique generally outperforms the MK-Means and MMD approaches. Overall, these results provide strong evidence that we can accurately detect concept drift in malware classification models.

CommentsTo appear as a chapter in the book "Artificial Intelligence for Cyber Defense in Emerging Threats", to be published by Springer by early 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑