arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

减轻具有层次结构和依赖结构的数据中类别不平衡的影响

Mitigating The Effect of Class Imbalance in Data with Hierarchical and Dependable Structure

Bipin Chhetri, Deepika Giri, Avishek Kadel, Rabin Kumar Karki, Akbar Siami Namin

arXiv 2607.11994首次发表:更新:

发表机构

Texas Tech University; Cumberland University; Yeshiva University; University of Cumberlands(德克萨斯理工大学; 坎伯兰大学; 叶史瓦大学; 坎伯兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对网络安全漏洞分类中类别不平衡及层次依赖问题挑战,提出层次感知RoBERTa框架,通过可学习父类嵌入纳入CWE结构信息,实验证明其优于过采样方法,在CWE研究概念数据集上加权F1分数达0.76,少数类提升显著。

AI 中文摘要

使用通用弱点枚举(CWE)分类法对网络安全漏洞进行分类具有挑战性,因为类别极度不平衡且弱点类别之间存在强烈的层次依赖性。虽然过采样技术如合成少数过采样技术(SMOTE)和自适应合成采样(ADASYN)被广泛用于减轻类别不平衡,但它们对层次CWE文本分类的有效性仍未得到充分探索。本文提出了一个层次感知的RoBERTa框架,通过可学习的父类嵌入明确纳入CWE结构信息,保持分类一致性。实验表明,高维嵌入空间中的合成插值违反了CWE层次结构的固有父子约束,对经典机器学习模型益处不大,却会持续降低深度学习架构的性能。在CWE研究概念数据集上评估,该模型在无数据增强时加权F1分数达到0.76,优于所有基线,少数类有显著提升,如Class类别的F1分数从BERT基线的0.40提高到0.60。结果表明,层次感知表示学习是结构化漏洞分类中比过采样更有原则的替代方法。

英文摘要

Classifying cybersecurity vulnerabilities using the Common Weakness Enumeration (CWE) taxonomy is challenging due to extreme class imbalance and strong hierarchical dependencies among weakness categories. Although oversampling techniques such as Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) are widely adopted to mitigate class imbalance, their effectiveness for hierarchical CWE text classification remains largely unexplored. This paper proposes a Hierarchy-Aware RoBERTa framework that explicitly incorporates CWE structural information through learnable parent-class embeddings, preserving taxonomic consistency. Our experiments demonstrate that synthetic interpolation in high-dimensional embedding spaces violates the inherent parent-child constraints of the CWE hierarchy, offering only marginal benefits for classical ML models while consistently degrading deep learning architectures. Evaluated on a CWE Research Concept dataset, the proposed model achieves a weighted F1-score of 0.76 without data augmentation, outperforming all baselines with notable gains on minority classes, including the Class category whose F1-score improved from 0.40 to 0.60 over the BERT baseline. Our results suggest that hierarchy-aware representation learning is a more principled alternative to oversampling for structured vulnerability classification.

Comments8 pages, 2 figures, 4 tables; preprint submitted to IEEE COMPSAC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑