发表机构
Northwestern University; Intel Corporation(西北大学; 英特尔公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对联邦深度聚类中的隐私泄露与数据异构问题,提出Fed-BRDECS框架,通过局部样本稳定性损失和预测平衡采样及质心重启,在图像、文本及时间序列任务上优于现有基线。
AI 中文摘要
联邦深度聚类旨在从分散的非标注数据中学习有利于聚类的表示,同时保护客户端隐私。然而,深度嵌入聚类(DEC)风格的目标函数依赖于全局软分配统计信息,这要求客户端泄露其敏感信息。我们提出了Fed-BRDECS,一个隐私保护且异构感知的联邦深度嵌入聚类框架。Fed-BRDECS用局部可计算的样本稳定性损失替代全局归一化的聚类目标,避免了传输局部软分配分布。为应对非独立同分布(non-IID)的客户端分布,我们引入了预测平衡采样,该采样在不需真实标签的情况下对局部稀有的预测簇进行过采样,以及质心级重启,定期刷新有偏或失活的质心。在图像和文本聚类基准上的实验表明,Fed-BRDECS在IID和非IID划分下均持续优于代表性的联邦聚类和深度聚类基线。我们进一步展示了其在联邦时间序列异常检测中的适用性,在不增加推理时间成本的情况下提升了基于重建的检测器性能。
英文摘要
Federated deep clustering seeks to learn clustering-friendly representations from decentralized unlabeled data while preserving client privacy. However, Deep Embedded Clustering (DEC)-style objectives depend on global soft-assignment statistics that require clients to reveal their sensitive information. We propose Fed-BRDECS, a privacy-preserving and heterogeneity-aware federated deep embedded clustering framework. Fed-BRDECS replaces the globally normalized clustering objective with a locally computable sample-stability loss, avoiding the transmission of local soft-assignment distributions. To tackle non-IID client distributions, we introduce prediction-balanced sampling, which oversamples locally rare predicted clusters without requiring ground-truth labels, and centroid-level restarting, which periodically refreshes biased or inactive centroids. Experiments on image and text clustering benchmarks show that Fed-BRDECS consistently outperforms representative federated clustering and deep clustering baselines under both IID and non-IID partitions. We further demonstrate its applicability to federated time-series anomaly detection, where it improves reconstruction-based detectors without adding inference-time cost.