arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36762cs.LGcs.DCstat.ML

联邦聚类中的未知局部与全局簇数量

Federated Clustering with Unknown Local and Global Cluster Cardinalities

Mitushi Goyal, Tarun S., Riddhanya Senapathi, Arun Raman

首次发表
浏览论文内容

中文总结 AI 辅助

针对联邦聚类中客户端与服务器均未知簇数量的场景,提出两阶段框架:客户端用自适应分裂-合并(ASM)估计局部簇数,再由 FedGEM 聚合,实验表明其优于现有无标签方法。

中文摘要 AI 辅助

不需要全局簇数量 $K$ 的联邦聚类方法仍然假设每个客户端知道其局部数量 $K_g$。当客户端对其数据的了解程度不高于服务器时(例如在独立运营的工业站点中进行故障诊断),这一假设难以成立。我们提出一个两阶段框架,其中两种数量均未知:每个客户端首先从自身数据估计 $K_g$,然后一个需要局部数量的聚合器(如 FedGEM)使用这些估计值代替真实值。对于第一阶段,我们引入自适应分裂-合并(ASM),该方法通过基于 BIC 的分裂来增长球形高斯混合,然后合并多余的成分。ASM 不使用标签,仅在保留的客户端数据上选择其超参数,并且不假设簇如何在客户端之间共享。我们推导出一个闭式分裂准则,其临界簇大小随各向异性增加而减小,随维度增加而增大,并经验性地表明过碎片化随每个簇的点数增加而增长,而联邦将其在客户端之间分配。在八个数据集上,ASM 与 FedGEM 的平均 ARI 达到 0.333,而次优的无标签估计器为 0.256,提供真实局部数量时为 0.361。它还能给出最可靠的 $K$ 全局估计,并且在客户端大小与局部基数解耦时具有鲁棒性。

英文摘要

Federated clustering methods that do not require the global number of clusters $K$ still assume that each client knows its local number $K_g$. This assumption is hard to justify when clients know no more about their data than the server does, as in fault diagnosis across independently operated industrial sites. We propose a two-phase framework in which neither count is known: each client first estimates $K_g$ from its own data, and an aggregator that requires local counts, such as FedGEM, then uses these estimates in place of the true values. For the first phase we introduce Adaptive Split--Merge (ASM), which grows a spherical Gaussian mixture by BIC-driven splitting and then merges excess components. ASM uses no labels, selects its hyperparameters on held-out client data only, and makes no assumption about how clusters are shared across clients. We derive a closed-form split criterion whose critical cluster size falls with anisotropy and rises with dimension, and show empirically that over-fragmentation grows with the number of points per cluster, which federation divides among clients. Across eight datasets, ASM with FedGEM attains a mean ARI of 0.333, against 0.256 for the next best label-free estimator and 0.361 when the true local counts are supplied. It also gives the most reliable global estimates of $K$ and is robust when client size is decoupled from local cardinality.

发表机构

  • BITS Pilani K. K. Birla Goa Campus(比拉理工学院K. K. Birla果阿校区)

机构由 AI 辅助整理,请以论文原文为准。

↑