发表机构
North Carolina A&T State University; Tribhuvan University; University of Tennessee-Knoxville; University of South Dakota(北卡罗来纳农工州立大学; 特里布文大学; 田纳西大学诺克斯维尔分校; 南达科他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对胸部X射线分类任务,在四大洲的四个队列上,通过联邦LoRA适配BiomedCLIP,提升了模型性能,且FlexLoRA的SVD乘积空间聚合是关键,无需集中数据即可实现协同适配。
AI 中文摘要
联邦学习(FL)允许各机构在不交换数据的情况下训练共享模型,而低秩适配(LoRA)仅通过传输紧凑的低秩更新即可实现大规模应用。生物医学成像非常适合这种组合场景:患者数据受隐私法规限制,且各机构的扫描仪、协议和计算资源差异巨大。这种异质性引发了一个问题:应如何聚合联邦LoRA更新,随着多模态视觉语言模型成为医学图像分析的核心,该问题愈发紧迫。我们对BiomedCLIP进行联邦参数高效微调(PEFT),以在三大洲(美国、越南、西班牙)的四个公共队列上进行胸部X射线分类任务的基准测试。联邦LoRA适配使所有四个队列的共享类别AUC从0.687提升至0.802,表明该提升源于联邦适配而非预训练模型的零样本能力。相较于孤立的单队列训练,联邦学习可改善较弱队列的性能,同时基本保留最强队列的性能,并接近汇集所有数据的集中式基准(AUC为0.812)。FlexLoRA引入的基于奇异值分解(SVD)的乘积空间聚合对该提升至关重要(朴素因子平均会使平均AUC下降0.097),而在我们的单种子运行中,带漂移校正的优化器(FedProx)相比FedAvg无明显优势,这与LoRA的低秩更新已限制客户端漂移的结论一致。因此,生物医学视觉语言模型可在不集中数据的情况下,跨异构、地理分布的机构进行协同适配。
英文摘要
Federated learning (FL) lets institutions train a shared model without exchanging data, and Low-Rank Adaptation (LoRA) makes this practical at scale by communicating only compact low-rank updates. Biomedical imaging is a compelling setting for this combination: patient data are archived behind privacy regulations, and institutions differ widely in scanners, protocols, and compute. Such heterogeneity raises the question of how federated LoRA updates should be aggregated, increasingly pressing as multimodal vision-language models become central to medical image analysis. We benchmark federated Parameter-efficient fine-tuning (PEFT) of BiomedCLIP for chest radiograph classification across four public cohorts on three continents (USA, Vietnam, Spain). Federated LoRA adaptation improves shared-class AUC on all four cohorts over the unadapted BiomedCLIP backbone (mean 0.687 to 0.802), showing that the gains come from federated adaptation rather than from the pretrained model's zero-shot ability. Relative to isolated single-cohort training, federation improves the weaker cohorts while largely preserving the strongest and approaches a centralized reference (0.812) that pools all data. The singular value decomposition (SVD)-based product-space aggregation introduced by FlexLoRA is essential to this gain (naive factor averaging drops mean AUC by 0.097), whereas a drift-correcting optimizer (FedProx) shows no benefit over FedAvg in our single-seed runs, consistent with LoRA's low-rank updates already limiting client drift. Biomedical vision-language models can thus be adapted collaboratively across heterogeneous, geographically distributed institutions without centralizing data.