发表机构
Faculty of Business and Commerce, Kansai University(关西大学商学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建含11种增强方法的基准,在7个数据集上测试后发现,基于LLM的增强性能不优于经典的EmbSMOTE,类结构保真度而非表面多样性是关键,建议检索式过采样作为不平衡多分类默认方法。
AI 中文摘要
随着大型语言模型(LLMs)的快速发展,生成式数据增强在自然语言处理的不平衡文本分类中受到了大量关注。然而,迄今为止,还没有实证基准将基于LLM的增强方法与嵌入空间SMOTE式检索(EmbSMOTE)进行比较,而EmbSMOTE是不平衡分类的强大经典参考方法。在本研究中,研究人员新构建了一个包含11种增强方法的受控基准,这些方法涵盖经典扰动、嵌入空间检索和基于LLM的生成,在7个公共文本分类数据集上进行测试,这些数据集的类别数K为2至28,不平衡比例从1.1到超过500;每个单元使用5个随机种子,通过宏F1、Welch t检验、5个分布度量以及基于Qwen3-8B的LLM家族敏感性分析进行评估。实验结果表明,所有基于LLM的方法在统计上与EmbSMOTE相当或更差,且随着不平衡程度增加,性能差距单调扩大,在GoEmotions-28上的宏F1差值ΔF1_macro约为0.063。此外,研究人员观察到,表面层面的独特性与下游性能几乎没有相关性,而LLM特有的人工产物(如文本拉长和标签分布均匀化)与分类准确率呈负相关。与6种基于LLM的和4种经典增强基线相比,这些结果表明,有效的变量不是表面层面的多样性,而是类条件结构保真度,即增强样本保留训练分布的类条件几何结构的程度。因此,检索式过采样应作为不平衡多分类的默认方法,而基于LLM的增强在实际部署前应设定更高的实证标准。
英文摘要
With the rapid advancement of large language models (LLMs), generative data augmentation has attracted considerable attention for imbalanced text classification in natural language processing. However, no empirical benchmark to date has compared LLM-based augmentation against the embedding-space SMOTE-style retrieval (EmbSMOTE), a strong classical reference for imbalanced classification. In this study, a controlled benchmark of 11 augmentation methods, spanning classical perturbation, embedding-space retrieval, and LLM-based generation, is newly constructed on seven public text classification datasets covering class counts $K=2$-$28$ and imbalance ratios of 1.1 to over 500, evaluated with five random seeds per cell using macro F1, Welch's $t$-tests, five distributional metrics, and an LLM-family sensitivity analysis based on Qwen3-8B. The experimental results reveal that all LLM-based methods are statistically equivalent or inferior to EmbSMOTE, with the performance gap widening monotonically as imbalance increases and reaching $Δ\text{F1}_\text{macro}\!\approx\!0.063$ on GoEmotions-28. Furthermore, it is observed that surface-level uniqueness has negligible correlation with downstream performance, whereas LLM-specific artifacts, such as text elongation and label-distribution uniformization, are negatively associated with classification accuracy. Compared with six LLM-based and four classical augmentation baselines, these results demonstrate that the effective variable is not surface-level diversity but class-conditional structural fidelity, namely the degree to which augmented samples preserve the class-conditioned geometry of the training distribution. Accordingly, retrieval-based oversampling should be adopted as the default for imbalanced multi-class classification, and a higher empirical bar should be required before LLM-based augmentation is deployed in practice.