arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11309stat.ME

部分损坏下有限混合模型的拜占庭容错分布式学习

Byzantine-tolerant distributed learning of finite mixture models under partial corruptions

  • Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院)
  • Department of Statistics, University of British Columbia(不列颠哥伦比亚大学统计学系)
  • School of Statistics, East China Normal University(华东师范大学统计学院)
  • Key Laboratory of Advanced Theory and Application in Statistics and Data Science-MOE, East China Normal University(华东师范大学统计与数据科学高级理论与应用教育部重点实验室)
  • Department of Statistics and Data Science, National University of Singapore(新加坡国立大学统计与数据科学系)

机构由 AI 辅助整理,请以论文原文为准。

Yimei Zhang, Jiahua Chen, Xiaozhou Wang, Yan Shuo Tan, Qiong Zhang

AI总结:

针对分布式有限混合模型估计中分量级拜占庭损坏问题,提出分量级过滤混合约简(CFMR)方法,通过数据驱动锚点、对齐过滤和混合约简聚合,达到预言机收敛速率,实验验证其有效性。

AI中文摘要:

有限混合模型刻画了异质性总体,并越来越多地通过分裂-征服程序拟合到分布式数据上,该程序在中心服务器上聚合局部混合估计。当传输的局部混合估计被部分或完全损坏时,聚合步骤可能受到严重损害。为防范拜占庭故障,现有鲁棒聚合方法针对局部混合估计要么完全真实、要么完全损坏的情形而开发。当仅部分分量估计被损坏时,此类方法可能丢弃有用信息。我们考虑分量级拜占庭故障,即某些分量估计可能被损坏,但对于每个混合分量,相应局部估计的大多数仍保持真实。我们提出分量级过滤混合约简(CFMR),它选择数据驱动的锚点,对齐传输的分量,通过多数半径过滤每个对齐的簇,并通过混合约简聚合保留的估计。通过仅过滤不可靠的分量,CFMR保留了来自部分损坏机器的真实信息,而无需知道故障率。我们为CFMR建立了自适应收敛界,并表明在适当条件下,它达到了若真实分量估计事先已知所能实现的预言机速率。模拟和真实数据应用表明,CFMR保持接近分量级预言机,而整机过滤和无保护聚合可能大幅恶化。

英文摘要:

Finite mixture models characterize heterogeneous populations and are increasingly fitted to distributed data using split-and-conquer procedures that aggregate local mixture estimates at a central server. The aggregation step can be seriously compromised when transmitted local mixture estimates are partially or completely corrupted. To guard against Byzantine failures, existing robust aggregation methods have been developed for settings in which a local mixture estimate is either entirely authentic or entirely corrupted. Such methods can discard useful information when only some component estimates are corrupted. We consider component-wise Byzantine failure, in which some component estimates may be corrupted, but for each mixture component, a majority of the corresponding local estimates remain authentic. We propose component-wise filtered mixture reduction (CFMR), which selects a data-driven anchor, aligns transmitted components, filters each aligned cluster by a majority radius, and aggregates the retained estimates through mixture reduction. By filtering out only unreliable components, CFMR preserves authentic information from partially corrupted machines without requiring knowledge of the failure rates. We establish an adaptive convergence bound for CFMR and show that, under suitable conditions, it attains the oracle rate that would be achieved if the authentic component estimates were known in advance. Simulations and a real-data application show that CFMR remains close to the component-level oracle, whereas whole-machine filtering and unprotected aggregation can deteriorate substantially.

↑