arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于元数据的分布式期望最大化的加密兼容聚类联邦学习

Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata

Michael Ben Ali, Imen Megdiche, André Péninou, Olivier Teste

arXiv 2607.28338首次发表:更新:

发表机构

UT3; IRIT; CNRS; INU Champollion; ISIS; UT2J(图卢兹第三大学; 信息论与电信研究所; 法国国家科学研究中心; 尚波利翁国立大学学院; 信息系统与安全研究所; 图卢兹第二让·饶勒斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FLAMECHE将基于元数据的CFL转化为分布式EM过程,实现加密兼容的聚类,优化了CFL三难问题的性能权衡,提升了客户端模型有效性。

AI 中文摘要

聚类联邦学习(CFL)通过将具有相似数据分布的客户端分组以实现有效训练,解决联邦环境中的数据异质性问题。现有方法在隐私保护、通信成本和计算效率之间存在权衡,我们将其形式化为CFL三难问题,即提升其中两个维度会以牺牲第三个维度为代价。一种主流范式依赖元数据(即客户端数据集与服务器共享的低维表示)实现通信和计算高效的聚类,但这类方法与标准FL隐私保护机制不兼容。为解决此限制,我们提出FLAMECHE,其将基于元数据的CFL重新表述为分布式期望最大化(EM)过程,将服务器更新限制为加法操作同时保持效率,该设计使其与实用的安全FL方案兼容。我们在多个异构场景下的多个数据集上进行了大量实验,结果表明FLAMECHE提升了客户端模型的有效性,实现了加密兼容的基于元数据的聚类,优化了其在CFL三难问题中的定位。

英文摘要

Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation, communication cost, and computational efficiency. We formalize this as the CFL trilemma, according to which improving two of these dimensions comes at the expense of the third. A prominent paradigm relies on metadata (i.e., low-dimensional representations of client datasets shared with the server) to enable communication- and computation-efficient clustering. However, such approaches are not compatible with standard FL privacy-preserving mechanisms. To address this limitation, we propose FLAMECHE, which reformulates metadata-based CFL as a distributed Expectation-Maximization (EM) procedure, restricting server updates to additive operations while preserving efficiency. This design enables compatibility with practical secure FL schemes. We conducted extensive experiments on multiple datasets under various heterogeneous scenarios. Results show that FLAMECHE improves the effectiveness of client models. It enables encryption-compatible metadata-based clustering, enhancing its positioning within the CFL trilemma.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑