arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAT-GS:通过校准门控与融合操作实现平衡多模态学习

CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery

Mahir Shahriar Tamim, Sharjil Khan, Md. Samiul Alim, Tanvir Ahmed Khan, Shafin Rahman, Nabeel Mohammed

arXiv 2608.24947首次发表:更新:

发表机构

North South University(北南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CAT-GS是一种无需修改模型的优化控制器,通过校准门控与融合操作解决多模态学习的三种失效模式,在多类基准上提升或匹配多模态准确率,且门控更平稳、融合冲突更少。

AI 中文摘要

端到端训练多模态神经网络时,常出现由三种耦合失效模式导致的不稳定神经动力学,降低学习效果:(i)模态失衡,即某一分支主导基于梯度的优化;(ii)门控不稳定,即带噪的置信度线索引发不稳定的模态选择;(iii)融合干扰,即特定模态的梯度在共享融合层发生冲突。我们提出CAT-GS(Calibrated, Adaptive, Thresholded Gating with Fusion Surgery,带融合操作的校准自适应阈值门控),这是一种面向智能计算应用的基于神经动力学的优化控制器。CAT-GS在反向传播阶段运行,无需修改模型架构、融合模块或任务损失。通过温度缩放与指数移动平均(EMA)平滑校准教师模型衍生的可靠性,CAT-GS采用边际阈值策略在预热丢弃、弱模态优先和弱偏向融合间切换,通过 capped 梯度预算重归一化在激进门控下稳定梯度幅度,并在主要共享瓶颈处应用仅融合的PCGrad以减少破坏性跨模态干扰。我们在音频-视觉多模态模式识别基准(CREMA-D、AV-MNIST和VGGSound)、三模态设置(UR-FUNNY)、受控合成数据(CG-MNIST)及额外跨域基准(AVE和CMU-MOSI)上评估CAT-GS。CAT-GS在各设置下均优于或匹配强大的感知失衡基线(包括OGM-GE、G²D和UMT)的融合多模态准确率,且门控行为更平稳,融合冲突梯度更少。

英文摘要

End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-specific gradients conflict at the shared fusion layer. We propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing applications. CAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses. Through calibration of teacher-derived reliability via temperature scaling and EMA smoothing, CAT-GS stabilizes neural dynamics using a margin-thresholded policy to switch between warm-up dropout, weak-modality prioritization, and weak-biased blending, stabilizes gradient magnitudes under aggressive gating via capped gradient-budget renormalization, and applies fusion-only PCGrad to reduce destructive cross-modal interference at the primary shared bottleneck. We evaluate CAT-GS on audio--visual multimodal pattern recognition benchmarks (CREMA-D, AV-MNIST, and VGGSound), a tri-modal setting (UR-FUNNY), controlled synthetic data (CG-MNIST), and additional cross-domain benchmarks (AVE and CMU-MOSI). CAT-GS improves or matches fused multimodal accuracy against strong imbalance-aware baselines (including OGM-GE, G$^2$D, and UMT) across settings, and yields smoother gating behavior with fewer conflicting fusion gradients.

CommentsThis article is accepted in Neurocomputing Journal

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑