面向多模态学习的自适应层级表示联盟
Adaptive Hierarchical Representation Alliance for Multimodal Learning
浏览论文内容
中文总结 AI 辅助
该研究针对多模态模型的语义粒度不匹配问题,提出自适应层级表示联盟框架,经六个基准实验验证,其性能优于强基线且在噪声、缺失模态场景下更鲁棒。
中文摘要 AI 辅助
多模态模型通常将语言、视觉和音频对齐到单一的最终层隐空间,隐含假设任务相关证据在各模态中出现于相同语义深度。通过逐层 CKA 分析,我们发现该假设会导致语义粒度不匹配:文本线索通常需要更深的上下文抽象,而视觉和声学线索往往在浅层或中层提供有区分度的感知证据。这种不匹配会抹平细粒度的模态私有线索,降低在有噪声、不平衡或缺失输入下的可靠性。为解决此问题,我们提出自适应层级表示联盟(Adaptive Hierarchical Representation Alliance, AHRA),这是一种层级共享-私有专家框架。AHRA 将每个模态跨语义层级分解为共享流和私有流,通过共享对齐和私有去相关进行正则化,通过跨模态专家路由共享信息,并在稀疏度控制的软门控机制(前景检查)引导下,用模态特定专家增强任务相关的私有 token。层级协同融合模块随后执行层内专家协调和层间语义选择。在图像-文本分类、多模态意图识别和三模态情感分析的六个基准上的实验表明,AHRA 始终优于强基线,且在有噪声和缺失模态的设置下保持鲁棒性。
英文摘要
Multimodal models often align language, vision, and audio in a single final-layer latent space, implicitly assuming that task-relevant evidence emerges at the same semantic depth across modalities. Using layer-wise CKA analysis, we observe that this assumption leads to semantic granularity mismatch: textual cues usually require deeper contextual abstraction, whereas visual and acoustic cues often provide discriminative perceptual evidence in shallow or middle layers. This mismatch can flatten fine-grained modality-private cues and reduce reliability under noisy, imbalanced, or missing inputs. To address this, we proposed Adaptive Hierarchical Representation Alliance (AHRA), a hierarchical shared--private expert framework. AHRA factorizes each modality into shared and private streams across semantic levels, regularizes them with shared alignment and private decorrelation, routes shared information through a cross-modal expert, and enhances task-relevant private tokens with modality-specific experts guided by a sparsity-controlled soft-gating mechanism (foreground exam). A hierarchical co-fusion module then performs intra-level expert coordination and inter-level semantic selection. Experiments on six benchmarks across image-text classification, multimodal intent recognition, and trimodal sentiment analysis show that AHRA consistently improves over strong baselines and remains robust under noisy and missing-modality settings.
发表机构
- Fudan University(复旦大学)
- University of Southern California(南加州大学)
- Cornell University(康奈尔大学)
- University of New South Wales(新南威尔士大学)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。