arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

伪标签在半监督Android恶意软件归属中的分类器依赖性收益

Classifier-Dependent Benefits of Pseudo-Labeling for Semi-Supervised Android Malware Attribution

Md Rafid Islam, Zahid Hasan, Hafiz Abdur Rahman

arXiv 2609.29564首次发表:更新:

发表机构

North South University(南北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了六种分类器在CICMalDroid 2020数据集上的伪标签半监督学习效果,发现收益强依赖分类器类型,SVM增益最大,随机森林在低标注比例下受损,并指出约10%标注样本即可达到近最优性能。

AI 中文摘要

由于特征维度高、类别不平衡以及专家标注数据成本高昂,检测和分类Android恶意软件家族仍然具有挑战性。半监督学习(SSL)提供了一种利用未标注样本的方法,但先前的研究很少测试SSL的收益是否在不同分类器类型间具有泛化性,或报告统计显著性。我们在CICMalDroid 2020数据集上,对六种分类器(LightGBM、XGBoost、随机森林、逻辑回归、MLP和SVM)进行了伪标签的系统性评估,采用五折分层交叉验证,并在五种标注比例(1%-20%)下进行了配对t检验。我们发现SSL的收益强烈依赖于分类器:SVM显示出最大的显著增益(在5%标注时准确率提升+4.4%,p=0.0028),LightGBM有适度提升(在2%-5%标注时提升+0.8%至+1.3%),而随机森林在低标注比例下受到显著损害(在1%标注时准确率下降-3.1%)。逐类分析显示,SSL不成比例地惠及最难分类的家族,其中Adware的F1分数提升了+13.8个百分点,而已被良好分类的Benign类仅提升+0.8。我们进一步表明,大约800个标注样本(数据集的10%)在所有分类器上都能产生接近最优的性能。这些发现为在Android恶意软件分类中何时以及使用哪种分类器进行伪标签提供了实用指导。

英文摘要

Detecting and classifying Android malware families remains challenging due to high feature dimensionality, class imbalance, and the high cost of expert-labeled data. Semi-supervised learning (SSL) offers a way to leverage unlabeled samples, but prior works rarely test whether SSL benefits generalize across classifier types or report statistical significance. We present a systematic evaluation of pseudo-labeling across six classifiers (LightGBM, XGBoost, Random Forest, Logistic Regression, MLP, and SVM) on the CICMalDroid 2020 dataset, using five-fold stratified cross-validation and paired t-tests across five labeled ratios (1-20%). We find that SSL benefit is strongly classifier-dependent: SVM shows the largest significant gain (+4.4% accuracy at 5% labels, p = 0.0028), LightGBM improves modestly (+0.8 to +1.3% at 2-5% labels), while Random Forest is significantly harmed at low label ratios (-3.1% at 1% labels). Per-class analysis reveals SSL disproportionately benefits the hardest-to-classify families, with Adware F1 improving by +13.8 percentage points versus only +0.8 for the already well-classified Benign class. We further show that approximately 800 labeled samples (10% of the dataset) yield near-optimal performance across all classifiers. These findings offer practical guidance on when and with which classifier pseudo-labeling is worthwhile for Android malware classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑