arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21664cs.LGstat.ML

基于测度量化的多域聚类

Multi-Domain Clustering via Measure Quantization

Rafael Pereira Eufrazio, Eduardo Fernandes Montesuma, Charles Casimiro Cavalcante

首次发表
浏览论文内容

中文总结 AI 辅助

提出基于测度量化的多域聚类框架,通过最小化概率度量学习共享原型,结合最优传输分配,在小批量优化下高效且优于基线。

中文摘要 AI 辅助

聚类是数据分析中的一项基本任务,通常通过基于质心的方法(如K均值)来解决。在这项工作中,我们提出了一个基于测度量化的多域聚类通用框架:给定来自多个域的样本,我们通过最小化每个域的概率测度与原型测度之间的概率度量(如Sinkhorn散度或最大均值差异)来学习一组共享的聚类原型。然后,数据点通过最近质心或最优传输(一种耦合域内所有样本的协作策略)被分配到聚类中。小批量优化策略使拟合和分配都具有可扩展性,在保持聚类性能的同时降低了内存和计算成本。在涵盖图像、音频和传感器数据的5个多域基准上的实验结果表明,我们基于Sinkhorn的方法始终优于经典和多域聚类基线,并且这种优势在扩展到数十万样本时仍然保持。

英文摘要

Clustering is a fundamental task in data analysis, typically addressed through centroid-based methods such as K-means. In this work, we present a general framework for multi-domain clustering via measure quantization: given samples from multiple domains, we learn a shared set of cluster prototypes by minimizing a probability metric, such as the Sinkhorn divergence or the Maximum Mean Discrepancy, between each domain's probability measure and the measure of prototypes. Data points are then assigned to clusters either via nearest centroid, or via optimal transport, a collaborative strategy that couples all samples within a domain. A mini-batch optimization strategy makes both fitting and assignment scalable, reducing memory and computational cost while preserving clustering performance. Experimental results on 5 multi-domain benchmarks spanning image, audio and sensor data show that our Sinkhorn-based method consistently outperforms classical and multi-domain clustering baselines, and that this advantage persists when scaling to hundreds of thousands of samples.

发表机构

  • Instituto Federal de Educação, Ciência e Tecnologia do Ceará(塞阿拉联邦教育与科学技术学院)
  • Federal University of Ceara(塞阿拉联邦大学)
  • Sigma Nova Science(Sigma Nova Science公司)

机构由 AI 辅助整理,请以论文原文为准。

↑