arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TEAMMix:用于大语言模型增强弱监督分层文本分类的类别体系丰富增强与少数类增强混合策略

TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

Jian Zhang, Zhuohao Yang, Songlin Lei, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang

arXiv 2608.11044首次发表:更新:

发表机构

School of Computer Science and Technology, Zhejiang University; ZJU-UIUC Institute, Zhejiang University; Shaoxing K3i Technology Co. Ltd; State Key Laboratory of CAD&CG, Zhejiang University(浙江大学计算机科学与技术学院; 浙江大学ZJU-UIUC研究院; 绍兴K3i科技有限公司; 浙江大学CAD&CG国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM应用于分层文本分类的局限,本文提出TEAMMix框架,通过丰富标签层次、生成伪样本并优化其质量,提升了细粒度及不平衡数据集上的分类性能。

AI 中文摘要

分层文本分类(HTC)是一项关键的文本挖掘任务,面临着标签层次结构复杂、类别不平衡等挑战。现有基于大语言模型(LLM)的方法因提示冗长、标签结构信息丢失等问题,难以高效应用于该任务。为解决这些局限,本文提出一种由LLM数据增强增强的弱监督HTC框架。该框架首先通过关键词生成与语料挖掘在语义层面丰富标签层次结构,提升模型对标签的理解;随后引导LLM生成伪样本以缓解长尾问题,并采用高斯混合模型进行基于置信度的重采样,优化生成数据的质量。实验结果表明,所提方法可有效提升LLM生成伪标签的可靠性,并在细粒度及不平衡数据集上显著增强分类性能。

英文摘要

Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing methods based on large language models (LLMs) struggle to be efficiently applied to this task due to issues like lengthy prompts and loss of label structural information. To address these limitations, this paper proposes a weakly supervised HTC framework enhanced by LLM-based data augmentation. The framework first enriches the label hierarchy semantically through keyword generation and corpus mining, thereby enhancing the model's understanding of labels. Subsequently, it guides the LLM to generate pseudo-samples to mitigate the long-tail problem, and employs a Gaussian mixture model for confidence-based resampling to optimize the quality of generated data. Experimental results demonstrate that the proposed method effectively improves the reliability of LLM-generated pseudo-labels and significantly enhances classification performance on fine-grained and imbalanced datasets.

CommentsAccepted by IEEE CSCWD 2026

DOI:10.1109/CSCWD68734.2026.11582679

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑