发表机构
School of Computer Science and Engineering, University of Electronic Science and Technology of China; Tianyijiaotong Technology Ltd.; School of Computer Science and Technology, Hainan University(电子科技大学计算机科学与工程学院; 天翼交通科技有限公司; 海南大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态学习中静态对齐与置信度融合的缺陷,提出ACML框架,通过动态三元组对齐和差异感知注意力校准,在多个基准数据集上取得优于现有方法的性能与鲁棒性。
AI 中文摘要
动态多模态学习旨在通过自适应地建模模态间的信息差异来学习鲁棒的表示。然而,现有方法仍存在两个局限性:(i)静态的跨模态对齐策略通常对所有样本施加统一约束,而忽略了样本间的差异,可能导致不合理的过度对齐;(ii)基于置信度或不确定性的融合方法往往未能充分考虑模态间的特征幅度和置信度差异。对于特征幅度差异显著或置信度差距较小的模态对,严格根据置信度对齐融合权重可能不可靠。为解决这些问题,我们提出了一种对齐与校准驱动的多模态学习框架(ACML)。具体而言,ACML包含一个动态跨模态三元组对齐模块,该模块对高置信度的正样本对强制执行强语义一致性,同时根据置信度差距鼓励高置信度与低置信度正样本对之间的多样化表示学习。此外,ACML引入了一种差异感知的注意力校准策略,该策略根据模态间的特征幅度和置信度差异自适应地调整注意力正则化,从而减轻由不合理融合约束引起的偏差。在多个多模态基准数据集上的大量实验表明,ACML在性能和鲁棒性上 consistently 优于近期最先进的方法。
英文摘要
Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
Comments17 pages