arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15687cs.LG

迈向联邦多模态图基础模型:一种拓扑感知多模态对齐框架

Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

  • Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Xunkai Li, Guohao Fu, Yuming Ai, Zhengyu Wu, Hongchao Qin, Rong-Hua Li, Guoren Wang

AI总结:

研究针对多模态属性图分散在不同隐私受限孤岛的问题,提出FedGAMMA框架,将联邦多模态图基础学习视为两阶段语义结构对齐问题,经实验验证该框架在下游任务中表现出色,超越诸多基线,在少样本学习场景下也有优势。

AI中文摘要:

多模态属性图(MAG)的节点除拓扑结构外还带有图像和文本等模态,广泛应用于社交平台、电子商务和生物医学网络等,比单模态图提供更丰富语义信号。实际中这些图分散在不同平台和机构的隐私受限孤岛中,学习可广泛转移的模型需要协作训练且不暴露原始数据,这处于多模态图学习和联邦学习的交叉点,但现有方法只涉及一方面。为应对这两个视角的挑战,我们提出FedGAMMA,将联邦多模态图基础学习视为联邦预训练和基于提示微调的两阶段语义结构对齐问题。预训练阶段,共享-私有语义增强器通过最优传输从模态特定信息中解耦跨模态共性,拓扑感知图融合模块通过语义残差图和双位置编码解耦语义和结构视图,双通道亲和力感知聚合机制从特征和图质心估计客户端相似度而不暴露原始数据。微调阶段,FedGAMMA通过轻量级图感知提示、具有可控探索的共享提示池和通道级提示同步来适应预训练编码器。在十二个多模态图数据集上的实验表明,FedGAMMA在下游任务中始终超过广泛的基线,增益高达12.96%。在少样本学习场景下,FedGAMMA在多域数据集的多个任务上进一步超越竞争基线,最高达5.71%。

英文摘要:

Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs. In practice, such graphs are fragmented across privacy-restricted silos owned by different platforms and institutions, so learning a broadly transferable model over them demands collaborative training that never exposes raw data. This places the task at the intersection of multimodal graph learning and federated learning, yet existing methods cover only one side of it. To address the challenges from these two perspectives, we propose FedGAMMA, casting federated multimodal graph foundation learning as a two-stage semantic-structural alignment problem of federated pre-training and prompt-based fine-tuning. During pre-training, a shared-private semantic enhancer disentangles cross-modal commonality from modality-specific information, aligning it through optimal transport, a topology-aware graph fusion module decouples semantic and structural views via semantic residual graphs and dual positional encodings, and a dual-channel affinity-aware aggregation mechanism estimates client similarity from feature and graph centroids without exposing raw data. During fine-tuning, FedGAMMA adapts the pretrained encoder through lightweight graph-aware prompts, a shared prompt pool with controlled exploration, and channel-wise prompt synchronization. Experiments on twelve multimodal graph datasets show FedGAMMA consistently surpassing a broad range of baselines across downstream tasks, with gains of up to 12.96%. FedGAMMA further outperforms competitive baselines accross multi-domain datasets on multiple tasks with up to 5.71% under few-shot learning scenario.

补充信息

↑