arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09240cs.LGcs.AI

双轴模态缺失下的多模态联邦学习

Multimodal Federated Learning under Dual-Axis Modality Missingness

Adiba Orzikulova, Jaehyun Kwak, Jaemin Shin, Yunqi Guo, Xiaomin Ouyang, Guoliang Xing, Steven Euijong Whang, Sung-Ju Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态联邦学习的双轴模态缺失问题,提出Flux框架,通过模态感知置信度调温与梯度解耦私有适配,在四个多模态数据集上取得最高平均宏F1,性能优于最强基准

中文摘要 AI 辅助

多模态联邦学习(FL)支持在对隐私敏感的健康感知与医疗场景中开展协作式建模,但实际部署中常出现双轴模态缺失:不同客户端拥有不同的模态集合,且单个样本可能仅包含本地可用模态的子集。现有方法通常分别处理这两个轴的问题。我们提出Flux,这是一个围绕两个互补组件构建的多模态联邦学习框架。第一,模态感知置信度调温(modality-aware confidence tempering)通过掩码感知的单模态监督为每个模态学习样本特定的置信度,并将来自观测模态的置信度估计融合为样本自适应的温度,该温度会根据证据质量与完整性调整预测的锐度。第二,梯度解耦的私有适配(gradient-decoupled private adaptation)仅将该温度应用于客户端私有预测通路,同时使用标准的未调温目标训练共享联邦模型。这使得在不允许依赖置信度的梯度干扰共享表示学习的情况下,实现样本特定的客户端本地置信度适配。在四个多模态数据集上,Flux在每个数据集上均达到最高的平均宏F1值,相较于最强的特定数据集基准,提升了0.8~2.2个点,平均提升1.6个点。额外分析表明,其具有良好的校准效果,温度对模态缺失与输入损坏均敏感,且在仅私有调温下共享优化更稳定。我们的代码可在此处获取:https://this.url

英文摘要

Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments often exhibit dual-axis modality missingness: clients have different modality sets, and individual samples may contain only subsets of the modalities available locally. Existing methods typically address these two axes separately. We propose Flux, a multimodal federated learning framework built around two complementary components. First, modality-aware confidence tempering learns sample-specific confidence for each modality through mask-aware unimodal supervision and fuses the confidence estimates from observed modalities into a sample-adaptive temperature that adjusts predictive sharpness according to evidence quality and completeness. Second, gradient-decoupled private adaptation applies this temperature only to a client-private prediction pathway, while training the shared federated model with a standard, untempered objective. This enables sample-specific, client-local confidence adaptation without allowing confidence-dependent gradients to perturb shared representation learning. Across four multimodal datasets, Flux achieves the highest average macro-F1 on every dataset, outperforming the strongest dataset-specific baseline by 0.8~2.2 points and by 1.6 points on average. Additional analyses demonstrate favorable calibration, temperature sensitivity to both modality missingness and input corruption, and more stable shared optimization under private-only tempering. Our code is available at https://github.com/AdibaOrz/Flux.

发表机构

  • KAIST(韩国科学技术院)
  • CUHK(香港中文大学)
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑