FedDRAW:面向异构多机构胸部X线片分类的联邦双声誉退火加权方法
FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification
浏览论文内容
中文总结 AI 辅助
针对异构多机构胸部X线片分类中联邦学习的样本规模偏见问题,提出FedDRAW方法,结合数据规模先验与参数相似度的双退火调度聚合权重,在CheXpert等数据集上较7种基线取得最优性能,可生成偏见更小的诊断模型。
中文摘要 AI 辅助
人工智能模型在医疗诊断中具有应用前景,但它们需要大量无偏数据,而医疗领域的数据分布在各医院,无法集中以保护患者隐私。联邦学习(FL)可解决该问题,各医院在本地训练共享诊断模型,患者数据保留在本地,训练在通信轮次中进行,每轮各医院在本地训练共享模型后返回服务器,服务器通过加权平均合并模型,聚合权重决定各机构知识对结果的影响。联邦平均(FedAvg)按本地样本数量设置权重,导致信息丰富但规模小的医院影响力始终较小,而信息不足的大规模客户端可能主导全局模型。我们提出联邦双声誉退火加权(FedDRAW),这是一种服务器端聚合方法,结合数据规模先验与客户端和全局参数在两个耦合退火调度下的余弦相似度。内部调度将客户端声誉从规模先验转向相似度,外部延迟退火调度在softmax逆温度上,使权重在早期和中期轮次保持选择性,收敛时放松至均匀。我们在两个胸部X线片数据集(CheXpert和ChestMNIST)的12种模拟客户端划分场景中评估FedDRAW,在相同本地训练设置下与7种联邦基线方法对比。FedDRAW在AUC及灵敏度与特异度的几何均值(GM)指标下,在所有8种方法中取得最高平均排名,经Friedman检验和Nemenyi事后分析确认,该差异具有统计学意义。调度两种信号而非仅按样本数量固定权重,可实现偏见更小的诊断模型。
英文摘要
Artificial intelligence models are promising for medical diagnosis, but they require large numbers of unbiased data, which in medicine are distributed across hospitals and cannot be centralized to protect patient privacy. Federated Learning (FL) addresses this, since hospitals train one shared diagnostic model while patient data remain local. Training proceeds in communication rounds, in which each hospital trains the shared model locally and returns it to the server for merging by weighted average. This aggregation weight determines whose institutional knowledge shapes the result. Federated averaging (FedAvg) sets it in proportion to local sample count, so a small but informative hospital is permanently assigned a small influence, andl argest clients could dominate the global model even when they are less informative. We propose Federated Dual Reputation Annealing Weighting (FedDRAW), a server-side aggregation method that combines a data-size prior with the cosine similarity between client and global parameters under two coupled annealing schedules. An inner schedule shifts client reputation from the size prior towards similarity. An outer, deferred annealing schedule on the softmax inverse temperature keeps the weighting selective in the early and middle rounds and relaxes it to uniformity at convergence. We evaluate FedDRAW on 12 simulated client-partition scenarios of two chest radiograph datasets (CheXpert and ChestMNIST), against seven federated baselines under identical local training settings. FedDRAW achieved the highest average rank among all eight methods under both AUC and the geometric mean (GM) of sensitivity and specificity, which a Friedman test with Nemenyi post-hoc analysis confirmed to be a statistically significant difference between the methods. Scheduling two signals, rather than fixing the weights by sample count alone, could enable less biased diagnostic models.
发表机构
- Institute for Predictive Deep Learning in Medicine and Healthcare, Justus-Liebig University(尤斯图斯-李比希大学医学与医疗保健预测深度学习研究所)
- Institute for Medical Informatics, University Medical Center Göttingen(哥廷根大学医学中心医学信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。