贝叶斯纵向联邦学习框架用于多元降秩高维回归
A Bayesian Vertical Federated Learning Framework for Multivariate Reduced-Rank High-Dimensional Regression
浏览论文内容
中文总结 AI 辅助
提出BayesVFLReg,一种用于多元高维降秩回归的贝叶斯纵向联邦学习框架,通过随机草图保护隐私,实现精确系数估计与特征选择,并具有理论保证和实证有效性。
中文摘要 AI 辅助
联邦学习(FL)已成为跨分散环境协作机器学习的主要隐私保护框架。尽管在水平联邦学习(HFL)方面已取得显著进展,其中具有共同特征的数据分布在各站点,但纵向联邦学习(VFL)中站点共享不同特征集的观测数据,仍较少被探索。推进用于VFL的贝叶斯高维多元降秩回归方法面临独特挑战:(a)严格的隐私法规阻止本地站点数据共享,以及(b)拟合局部回归忽略了诸如变量间相关性等关键建模方面。相比之下,HFL允许每个站点独立拟合可比模型。我们提出了一种新颖的贝叶斯VFL框架,用于多元高维降秩回归,称为BayesVFLReg,该框架在保护特征和响应隐私的同时实现精确的系数估计。参与站点使用共享的随机草图矩阵将局部变量压缩为隐私保护的草图。中央服务器收集这些草图,其中贝叶斯多元降秩回归使用高斯尺度混合先验。对于特征选择,我们引入了一种基于混合模型聚类的单步后处理策略,对绝对后验系数均值进行聚类,以区分每个响应变量的信号与噪声。BayesVFLReg对于大型高维数据集具有计算可扩展性,并促进高效的变量选择。理论上,我们建立了拟合密度落入以真实数据生成密度为中心的Hellinger球内的后验概率的尖锐非渐近界。比较模拟研究和真实数据分析表明,BayesVFLReg即使在特征相关的情况下也能可靠地识别稀疏特征效应。
英文摘要
Federated learning (FL) has emerged as a leading privacy-preserving framework for collaborative machine learning across decentralized environments. While considerable progress has been made in horizontal federated learning (HFL), where data with common features is distributed across sites, vertical federated learning (VFL), where sites share observations across distinct feature sets, remains less explored. Advancing Bayesian high-dimensional multivariate reduced-rank regression methods for VFL poses unique challenges: (a) stringent privacy regulations preventing local site data sharing, and (b) fitting local regressions overlooks essential modeling aspects like inter-variable correlations. In contrast HFL allows each site to fit a comparable model independently. We present a novel Bayesian VFL framework for multivariate high-dimensional reduced-rank regression, termed BayesVFLReg, which enables precise coefficient estimation while safeguarding both feature and response privacy. Participating sites use a shared random sketching matrix to compress local variables into privacy-preserving sketches. A central server collects these sketches where Bayesian multivariate reduced-rank regression uses Gaussian scale mixture priors. For feature selection, we introduce a single-step post-processing strategy based on mixture-model clustering of the absolute posterior coefficient means to distinguish signal from noise per response variable. BayesVFLReg is computationally scalable for large, high-dimensional datasets and facilitates efficient variable selection. Theoretically, we establish sharp non-asymptotic bounds on the posterior probability that the fitted density falls within a Hellinger ball centered at the true data-generating density. Comparative simulation studies and real-world data analyses show that BayesVFLReg reliably identifies sparse feature effects, even under feature correlation.
发表机构
- Texas A&M University(德克萨斯农工大学)
机构由 AI 辅助整理,请以论文原文为准。