发表机构
Institute of Medical Informatics, University of Münster(明斯特大学医学信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对群体学习中因特征异质性导致的问题,提出用随机森林解决,给出确定性和概率性推理时策略,在九个数据集上评估,结果显示该方法在多种场景下优于交集基线和本地训练模型。
AI 中文摘要
群体学习是一种去中心化协作学习机制,能让多个组织在无中央协调或直接数据共享的情况下训练共享模型。典型的水平群体学习通常假定各站点数据集共享相同特征集,但实际应用中各站点特征常部分重叠。这种特征异质性给随机森林等机器学习算法带来问题,决策树合并成全局随机森林时,若遍历遇到本地不可用特征的分割,推理会不明确。本文针对部分重叠特征空间下群体学习的特征异质性,提出多种确定性和概率性推理时策略,无需将训练限制在特征交集上。在九个数据集上评估方法,结果表明在多种场景下优于交集基线和本地训练模型。
英文摘要
Swarm Learning is a decentralized collaborative learning mechanism that allows multiple organizations to train a shared model without central coordination or direct data sharing. In typical horizontal Swarm Learning, datasets across sites are usually assumed to share the same feature set. However, in real-world applications, sites often have partially overlapping features because measurements, protocols, and available covariates differ across sites. This feature heterogeneity creates a practical issue for machine learning algorithms such as Random Forests. Specifically, when decision trees are pooled into a global Random Forest, inference at a given site can become ill-defined if a traversal encounters a split on a feature that is not available locally, often forcing organizations to discard site-specific variables upfront. In this paper, we address feature heterogeneity in Swarm Learning with Random Forests under partially overlapping feature spaces. We propose several deterministic and probabilistic inference-time strategies that resolve such missing splits without restricting training to the intersection of features. We evaluate the methods on nine datasets and demonstrate that they outperform both the intersection baseline and locally trained models across a broad range of scenarios.