arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于高斯混合模型(GMM)与大语言模型(LLM)的定向数据增强用于不平衡数据聚类

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

Noor Khalal, Abdallah Alaa-Eddine Djamai, Imed Keraghel, Mohamed Nadif

arXiv 2607.28635首次发表:更新:

发表机构

Centre Borelli UMR9010; Université Paris Cité; Kernix Software(博雷利中心UMR9010; 巴黎城市大学; Kernix软件公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无监督NLP任务中不平衡数据聚类难以捕捉少数主题的问题,本文提出整合GMM与LLM的定向数据增强方法,实验表明其可保持聚类性能并提升可解释性。

AI 中文摘要

在自然语言处理(NLP)领域,处理代表性不足的主题颇具挑战,尤其在聚类这类无监督任务中,模型可能无法充分捕捉少数主题。为解决该问题,本文提出一种新型无监督数据增强方法,整合高斯混合模型(GMMs)与大语言模型(LLMs)。凭借灵活性与鲁棒性,GMMs可识别数据中对应代表性不足区域的聚类,LLMs则生成合成文档以丰富这些聚类、提升其代表性。在多种不平衡文本数据集上开展的实验表明,本文方法在所有情况下均能保持聚类性能,且常可提升聚类可解释性,为改进无监督NLP任务中的数据表示提供了一种鲁棒且可扩展的解决方案。

英文摘要

In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not adequately capture minority topics. To tackle this challenge, our paper presents a novel unsupervised data augmentation method that integrates Gaussian Mixture Models (GMMs) and Large Language Models (LLMs). Due to their flexibility and robustness, GMMs can detect clusters corresponding to underrepresented areas in the data, while LLMs create synthetic documents to enrich these clusters and improve their representation. Experiments on various imbalanced text datasets demonstrate that our approach preserves clustering performance in all cases and often enhances cluster interpretability, offering a robust and scalable solution for improving data representation in unsupervised NLP tasks.

Journal refAdvances in Intelligent Data Analysis: 23rd International Symposium on Intelligent Data Analysis; IDA 2025; Proceedings; pp 246-260

DOI:10.1007/978-3-031-91398-3

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑