arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Fed-ReMasker:特征级缺失下的联邦表格数据填补

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou

arXiv 2609.28105首次发表:更新:

发表机构

ARTORG Center for Biomedical Engineering Research University of Bern; Graduate School for Cellular and Biomedical Sciences University of Bern(伯尔尼大学生物医学工程研究中心; 伯尔尼大学细胞与生物医学科学研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对联邦学习中特征级缺失问题,提出Fed-ReMasker方法,将掩码自编码器适配至联邦框架,利用跨中心知识填补未观测特征,在多种基准中显著优于现有方法,接近集中式模型性能。

AI 中文摘要

多中心临床研究和生物医学研究合作日益寻求利用各中心的数据来构建超越任何单一中心的泛化模型。这带来了两个不同的挑战:数据保护法规可能限制跨机构共享原始患者数据,而各中心可能在不同协议下仅收集部分重叠的特征集。联邦学习使得无需集中原始数据即可进行协作模型训练。然而,现有的联邦填补方法很少评估特征级缺失,即某些中心完全未观察到某些特征。为解决这一场景,我们将ReMasker掩码自编码器适配到联邦学习(Fed-ReMasker),使各中心能够利用跨协作中心学到的知识来填补本地从未观察到的特征。我们在一个基准测试中评估了Fed-ReMasker,该基准涵盖具有线性和非线性关系的合成数据集以及真实世界的表格数据集(包括临床数据)。该基准测试变化了中心数量、缺失比例和客户端异质性。在均匀基准中,Fed-ReMasker在93.2%的值级场景和96.7%的特征级场景中实现了最低的填补误差。它还能通过简单的联邦平均对客户端异质性保持鲁棒性,在所有36个值级场景中优于所有基线,在至少35/36个特征级场景中优于每个基线,并且平均比在汇总数据上训练的集中式模型仅差3.0%。

英文摘要

Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection regulations may restrict the sharing of raw patient data across institutions, while centers may collect only partially overlapping sets of features under different protocols. Federated learning enables collaborative model training without centralizing raw data. However, existing federated imputation methods rarely evaluate feature-level missingness, in which entire features are unobserved at some centers. To address this setting, we adapt the ReMasker masked autoencoder to federated learning (Fed-ReMasker), enabling centers to impute features never observed locally by leveraging knowledge learned across collaborating centers. We evaluate Fed-ReMasker in a benchmark spanning synthetic datasets with linear and nonlinear relationships and real-world tabular datasets, including clinical data. The benchmark varies the number of centers, the missingness ratios, and client heterogeneity. Fed-ReMasker achieves the lowest imputation error in 93.2% of value-level and 96.7% of feature-level scenarios in the homogeneous benchmark. It also remains robust to client heterogeneity using simple federated averaging, outperforming all baselines in all 36 value-level scenarios and each baseline in at least 35 of 36 feature-level scenarios, and comes within 3.0% on average of a centralized model trained on the pooled data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑