发表机构
Centre Borelli UMR9010; Université Paris Cité; Kernix Software(博雷利中心UMR9010; 巴黎城市大学; Kernix软件公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无监督NLP任务中不平衡数据聚类难以捕捉少数主题的问题,本文提出整合GMM与LLM的定向数据增强方法,实验表明其可保持聚类性能并提升可解释性。
AI 中文摘要
在自然语言处理(NLP)领域,处理代表性不足的主题颇具挑战,尤其在聚类这类无监督任务中,模型可能无法充分捕捉少数主题。为解决该问题,本文提出一种新型无监督数据增强方法,整合高斯混合模型(GMMs)与大语言模型(LLMs)。凭借灵活性与鲁棒性,GMMs可识别数据中对应代表性不足区域的聚类,LLMs则生成合成文档以丰富这些聚类、提升其代表性。在多种不平衡文本数据集上开展的实验表明,本文方法在所有情况下均能保持聚类性能,且常可提升聚类可解释性,为改进无监督NLP任务中的数据表示提供了一种鲁棒且可扩展的解决方案。
英文摘要
In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics. To tackle this challenge, our paper presents a novel unsupervised data augmentation method that integrates Gaussian Mixture Models (GMMs) and Large Language Models (LLMs). Due to their flexibility and robustness, GMMs can detect clusters corresponding to underrepresented areas in the data, while LLMs create synthetic documents to enrich these clusters and improve their representation. Experiments on various imbalanced text datasets demonstrate that our approach preserves clustering performance in all cases and often enhances cluster interpretability, offering a robust and scalable solution for improving data representation in unsupervised NLP tasks.
Journal refAdvances in Intelligent Data Analysis: 23rd International Symposium on Intelligent Data Analysis; IDA 2025; Proceedings; pp 246-260