arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向不平衡表格数据聚类的无监督深度学习集成方法

Ensemble of Unsupervised Deep Learning for Clustering Imbalanced Tabular Data

Pulock Das, Yina Hou, Md. Kamrozzaman Bhuiyan, Manar D. Samad

arXiv 2608.00346首次发表:更新:

发表机构

Tennessee State University; Enosis Solutions; North Carolina A&T State University(田纳西州立大学; 伊诺西斯解决方案公司; 北卡罗来纳农工州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对不平衡表格数据,提出两种无监督深度聚类集成方法,经16个二分类表格数据集实验,其在ACC、NMI、ARI指标上优于单个方法,可作为监督分类的有力替代方案。

AI 中文摘要

数据不平衡是监督分类中的重大挑战,多数类偏差会导致假阴性并高估分类准确率。无监督深度聚类因无需类别标签即可开展聚类表征学习,可免受类别不平衡的影响。深度聚类已被应用于图像、语言和图数据,但其在表格数据中的应用近年才兴起。本文率先研究了不同程度数据不平衡下,最先进深度聚类方法的性能。我们提出两种新型聚类集成方法:一种聚合不同嵌入维度下的深度聚类分配结果,另一种对表现最佳的聚类算法应用多数投票。在16个具有不同及人工诱导不平衡程度的二分类表格数据集上开展的实验,揭示了不同深度聚类方法的独特优势。平均而言,我们的集成方法在ACC、NMI和ARI指标上优于单个聚类方法,在无监督识别真实类别时,对数据不平衡具有更强的鲁棒性。因此,在不平衡数据场景中,深度聚类可作为监督分类的有力替代方案。

英文摘要

Data imbalance poses a major challenge in supervised classification, where the majority-class bias contributes to false negatives and overestimates classification accuracy. Unsupervised deep clustering can be immune to class imbalance because representation learning for clustering is performed without class labels. Deep clustering has been proposed for images, languages, and graphs, while its application to tabular data has only emerged recently. This paper is among the first to examine the performance of state-of-the-art deep clustering methods under varying levels of data imbalance. We introduce two novel cluster ensemble approaches: one aggregates deep clustering assignments across different embedding dimensions, and the other applies majority voting to the best-performing clustering algorithms. Experiments on 16 binary tabular datasets with varying and artificially induced levels of imbalance reveal distinct strengths of different deep clustering methods. On average, our ensemble methods outperform individual clustering methods in ACC, NMI, and ARI scores, offering greater resilience to data imbalance when identifying ground-truth classes without supervision. Therefore, in an imbalanced data scenario, deep clustering can serve as a strong alternative to supervised classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑