arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2411.17287cs.LG

面向小规模高维生物数据回归任务的隐私保护联邦无监督域自适应

Privacy-Preserving Federated Unsupervised Domain Adaptation for Regression on Small-Scale and High-Dimensional Biological Data

  • University of Tübingen(蒂宾根大学)

机构由 AI 辅助整理,请以论文原文为准。

Cem Ata Baykara, Ali Burak Ünal, Nico Pfeifer, Mete Akgün

更新

AI总结:

针对生物数据小规模、高维、分散及域偏移导致模型泛化难的问题,提出隐私保护联邦无监督域自适应回归方法freda,通过高斯过程联邦训练建模特征关系,在DNA甲基化数据年龄预测任务中性能接近集中式最优且保障隐私。

AI中文摘要:

机器学习模型常因数据收集差异和群体不同导致的域偏移,在小型异质数据集中泛化能力不足。这一挑战在生物数据中尤为突出,此类数据维度高、规模小且分散在各机构间。虽然联邦域自适应方法(FDA)旨在解决这些问题,但大多数现有方法依赖深度学习且聚焦分类任务,使其不适用于小规模高维应用场景。本研究提出freda,一种面向回归任务无监督域自适应的隐私保护联邦方法。与基于深度学习的FDA方法不同,freda是首个支持高斯过程联邦训练的方法,通过随机编码和安全聚合确保数据完全隐私,以建模复杂特征关系。这使其无需直接访问原始数据即可实现有效域自适应,适用于高维异质数据集应用。我们在DNA甲基化数据年龄预测这一具有挑战性的任务上评估freda,结果表明其性能可与集中式最先进方法媲美,同时保留完全数据隐私。

英文摘要:

Machine learning models often struggle with generalization in small, heterogeneous datasets due to domain shifts caused by variations in data collection and population differences. This challenge is particularly pronounced in biological data, where data is high-dimensional, small-scale, and decentralized across institutions. While federated domain adaptation methods (FDA) aim to address these challenges, most existing approaches rely on deep learning and focus on classification tasks, making them unsuitable for small-scale, high-dimensional applications. In this work, we propose freda, a privacy-preserving federated method for unsupervised domain adaptation in regression tasks. Unlike deep learning-based FDA approaches, freda is the first method to enable the federated training of Gaussian Processes to model complex feature relationships while ensuring complete data privacy through randomized encoding and secure aggregation. This allows for effective domain adaptation without direct access to raw data, making it well-suited for applications involving high-dimensional, heterogeneous datasets. We evaluate freda on the challenging task of age prediction from DNA methylation data, demonstrating that it achieves performance comparable to the centralized state-of-the-art method while preserving complete data privacy.

↑