arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22863cs.MM

面向多模态学习的自适应层级表示联盟

Adaptive Hierarchical Representation Alliance for Multimodal Learning

Chunlei Meng, Pengbin Feng, Jacqueline J. Pang, Chih-Ting Liao, Rong Fu, Zhaolu Kang, Zhongxue Gan, Chun Ouyang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多模态模型的语义粒度不匹配问题,提出自适应层级表示联盟框架,经六个基准实验验证,其性能优于强基线且在噪声、缺失模态场景下更鲁棒。

中文摘要 AI 辅助

多模态模型通常将语言、视觉和音频对齐到单一的最终层隐空间,隐含假设任务相关证据在各模态中出现于相同语义深度。通过逐层 CKA 分析,我们发现该假设会导致语义粒度不匹配:文本线索通常需要更深的上下文抽象,而视觉和声学线索往往在浅层或中层提供有区分度的感知证据。这种不匹配会抹平细粒度的模态私有线索,降低在有噪声、不平衡或缺失输入下的可靠性。为解决此问题,我们提出自适应层级表示联盟(Adaptive Hierarchical Representation Alliance, AHRA),这是一种层级共享-私有专家框架。AHRA 将每个模态跨语义层级分解为共享流和私有流,通过共享对齐和私有去相关进行正则化,通过跨模态专家路由共享信息,并在稀疏度控制的软门控机制(前景检查)引导下,用模态特定专家增强任务相关的私有 token。层级协同融合模块随后执行层内专家协调和层间语义选择。在图像-文本分类、多模态意图识别和三模态情感分析的六个基准上的实验表明,AHRA 始终优于强基线,且在有噪声和缺失模态的设置下保持鲁棒性。

英文摘要

Multimodal models often align language, vision, and audio in a single final-layer latent space, implicitly assuming that task-relevant evidence emerges at the same semantic depth across modalities. Using layer-wise CKA analysis, we observe that this assumption leads to semantic granularity mismatch: textual cues usually require deeper contextual abstraction, whereas visual and acoustic cues often provide discriminative perceptual evidence in shallow or middle layers. This mismatch can flatten fine-grained modality-private cues and reduce reliability under noisy, imbalanced, or missing inputs. To address this, we proposed Adaptive Hierarchical Representation Alliance (AHRA), a hierarchical shared--private expert framework. AHRA factorizes each modality into shared and private streams across semantic levels, regularizes them with shared alignment and private decorrelation, routes shared information through a cross-modal expert, and enhances task-relevant private tokens with modality-specific experts guided by a sparsity-controlled soft-gating mechanism (foreground exam). A hierarchical co-fusion module then performs intra-level expert coordination and inter-level semantic selection. Experiments on six benchmarks across image-text classification, multimodal intent recognition, and trimodal sentiment analysis show that AHRA consistently improves over strong baselines and remains robust under noisy and missing-modality settings.

发表机构

  • Fudan University(复旦大学)
  • University of Southern California(南加州大学)
  • Cornell University(康奈尔大学)
  • University of New South Wales(新南威尔士大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑