arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00623cs.LG

面向有效联邦多模态图学习的多层面异质性导航方法

Towards Effective Federated Multimodal Graph Learning via Navigating Multifaceted Heterogeneity

Yinlin Zhu, Di Wu, Yi Zhang, Xunkai Li, Wang Luo, Wei-Jin Huang, Miao Hu, Guocong Quan

首次发表
浏览论文内容

中文总结 AI 辅助

针对联邦多模态图学习的多层面异质性问题,本文提出FedTCR算法,通过两阶段范式与拓扑感知跨模态路由机制,在7个领域的图中心与模态中心任务上优于现有最优基线。

中文摘要 AI 辅助

多模态属性图(MAGs)是指节点承载跨多种模态的异质语义内容,边编码关系依赖的图结构,已在多个领域广泛应用。联邦多模态图学习(FMGL)将联邦图学习(FGL)扩展至MAGs,可在不暴露原始数据的前提下实现去中心化MAGs的协同优化。然而,直接将现有FGL方法应用于FMGL效果不佳,因为这些方法无法应对去中心化MAGs固有的多层面异质性,包括不同客户端目标带来的任务异质性、模态质量与语义领域差异导致的模态异质性,以及跨模态相关性低的不同拓扑模式产生的拓扑异质性。为解决这些挑战,本文提出首个针对FMGL的系统性算法——联邦拓扑感知跨模态路由多模态图学习(FedTCR)。对于任务异质性,FedTCR采用两阶段范式,包含联邦任务无关预训练与独立任务定向微调;为协同解决模态与拓扑异质性,FedTCR引入拓扑感知跨模态路由机制:每个客户端基于图结构,通过拓扑感知重要性加权聚合将模态特定知识提炼为紧凑原型;服务器随后评估这些结构感知原型间的跨客户端跨模态关系,将有信息的原型作为对比参考,驱动三级跨模态对比学习方案,在保留区分性的同时对齐跨客户端模态。在7个领域的实验表明,FedTCR在图中心与模态中心任务上均优于现有最优基线方法。

英文摘要

Multimodal-attributed graphs (MAGs), where nodes carry heterogeneous semantic content across multiple modalities while edges encode relational dependencies, have been widely adopted across diverse domains. Federated multimodal graph learning (FMGL) extends federated graph learning (FGL) to MAGs, enabling collaborative optimization across decentralized MAGs without exposing raw data. However, naively applying existing FGL methods to FMGL is insufficient, as they fail to navigate the multifaceted heterogeneity inherent in decentralized MAGs, including task heterogeneity across diverse client objectives, modality heterogeneity from discrepant modality quality and semantic domains, and topology heterogeneity arising from divergent topological patterns with low cross-modality correlation. To address these challenges, we propose Federated multimodal graph learning with Topology-aware Cross-modal Routing (FedTCR), the first systematic algorithm designed for FMGL. To handle task heterogeneity, FedTCR employs a two-stage paradigm that comprises federated task-agnostic pre-training followed by isolated task-oriented fine-tuning. To jointly address modality and topology heterogeneity, FedTCR introduces a topology-aware cross-modal routing mechanism. Concretely, each client distills modality-specific knowledge into compact prototypes via topology-aware importance-weighted aggregation informed by graph structure; the server then evaluates cross-client cross-modal relationships among these structure-informed prototypes and routes informative ones as contrastive references, driving a tri-level cross-modal contrastive learning scheme that jointly aligns cross-client modalities while preserving discrimination. Experiments across 7 domains demonstrate that FedTCR outperforms state-of-the-art baselines on both graph-centric and modality-centric tasks.

补充信息

↑