arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22835cs.CR

当标签噪声遇上类别不平衡:一种用于Android恶意软件家族分类的鲁棒框架

When Label Noise Meets Class Imbalance: A Robust Framework for Android Malware Family Classification

Haolan Zhang, Cuiying Gao, Fulin Zhao, Heng Li, Haoran Wang, Chang Luo, Tiejun Wu, Hui Shu, Wei Yuan

首次发表
浏览论文内容

中文总结 AI 辅助

针对Android恶意软件家族分类中标签噪声与类别不平衡的交互问题,提出RoMaC框架,采用自训练纠正噪声标签并区分头尾家族样本,结合类别重加权与多模型集成,在混淆场景下性能提升6%-20%。

中文摘要 AI 辅助

用于Android恶意软件家族分类的机器学习方法已取得较高准确率,但其应用受到两大挑战的阻碍。首先,广泛使用的代码混淆严重干扰了自动化标注过程,并在训练数据集中引入大量标签噪声。其次,训练数据集往往表现出严重的类别不平衡,导致家族分类模型性能不佳。尽管现有研究已针对标签噪声或类别不平衡提出了各种解决方案,但它们常常忽视这两个因素之间的相互作用。在类别不平衡条件下,难以学习的少数类样本的存在会显著削弱现有针对噪声样本对策的有效性。为联合处理标签噪声与类别不平衡,我们提出了一种鲁棒的Android恶意软件家族分类框架RoMaC。该框架采用自训练策略来纠正噪声标签,更重要的是,它对头部家族样本和尾部家族样本进行区别对待。这一设计有效缓解了类别不平衡对噪声鲁棒学习的不利影响。此外,RoMaC将类别重新加权机制与多模型集成学习相结合,从而提升了分类准确率和噪声鲁棒性。我们在由两个公开数据集构建的合并数据集上评估了RoMaC。当30%的样本被混淆时,RoMaC实现了0.803的总体Macro-F1分数和0.871的准确率,以及0.672的尾部类别Macro-F1分数和0.784的准确率。与现有方法相比,RoMaC在各种混淆场景和噪声水平下展现了6%-20%的性能提升。

英文摘要

Machine learning methods for Android malware family classification have achieved high accuracy, but their application is hindered by two major challenges. First, the widely used code obfuscation severely disrupts the automated labeling process and introduces substantial label noise into training datasets. Second, training datasets often exhibit severe class imbalance, leading to poor performance of family classification models. Although existing studies have proposed various solutions to either label noise or class imbalance, they often overlook the interplay between these two factors. Under class imbalance, the presence of hard-to-learn minority-class samples can significantly impair the effectiveness of existing countermeasures for noisy samples. To jointly address label noise and class imbalance, we propose a robust Android malware family classification framework, RoMaC. It employs a self-training strategy to correct noisy labels and, more importantly, discriminately treats head-family and tail-family samples. This design effectively mitigates the adverse impact of class imbalance on noise-robust learning. Moreover, RoMaC integrates a class reweighting mechanism with multi-model ensemble learning, thereby enhancing both classification accuracy and noise robustness. We evaluate RoMaC on a combined dataset constructed from two public datasets. When 30% of the samples are obfuscated, RoMaC achieves an overall Macro-F1 score of 0.803 and an accuracy of 0.871, as well as a tail-class Macro-F1 score of 0.672 and an accuracy of 0.784. Compared with existing methods, RoMaC demonstrates performance improvements of 6%-20% across various obfuscation scenarios and noise levels.

发表机构

  • Huazhong University of Science and Technology(华中科技大学)
  • The Hong Kong Polytechnic University(香港理工大学)
  • NSFOCUS Technologies Group Co., Ltd.(启明星辰信息技术集团股份有限公司)
  • Key Laboratory of Cyberspace Security, Ministry of Education(教育部网络空间安全学院)

机构由 AI 辅助整理,请以论文原文为准。

↑