AI 中文总结
针对多模态训练中的模态不平衡问题,提出方差校准动量VCMM,通过在线估计噪声与漂移并利用卡尔曼控制器自适应调整各模态动量,在四个基准上以低开销取得一致提升。
AI 中文摘要
多模态联合训练常常面临模态不平衡问题,其中主导模态会抑制其他模态的优化。现有方法主要通过调节梯度幅度或方向、修改优化目标或调整训练策略来平衡模态学习,且大多数干预措施集中于当前更新。然而,当与广泛使用的基于动量的优化器结合时,更新还包含来自先前梯度的累积信息,而仅靠当前步的调节无法显式处理这一点。为解决此问题,我们提出了方差校准动量(VCMM),该方法将梯度记忆适应于模态特定的梯度动态。具体而言,VCMM在线估计小批量噪声和时间漂移,并通过卡尔曼启发的控制器利用它们的相对强度来确定模态特定的动量。我们进一步对跨模态的控制信号进行中心化,并对时变一阶矩应用精确的偏差校正,从而无需额外的网络传递或显式学习率缩放即可实现自适应梯度记忆。在四个多模态基准上的实验表明,该方法在适度训练开销下取得了一致的改进。
英文摘要
Multimodal joint training often suffers from modality imbalance, where a dominant modality suppresses the optimization of others. Existing methods mainly balance modality learning by modulating gradient magnitudes or directions, modifying optimization objectives, or adjusting training strategies, with most interventions focusing on the current update. However, when combined with widely used momentum-based optimizers, the update also incorporates accumulated information from previous gradients, which is not explicitly addressed by current-step modulation alone. To address this issue, we propose Variance-Calibrated MomentuM (VCMM), which adapts gradient memory to modality-specific gradient dynamics. Specifically, VCMM estimates minibatch noise and temporal drift online and uses their relative strength to determine modality-specific momentum through a Kalman-inspired controller. We further center the control signal across modalities and apply exact bias correction for the time-varying first moment, enabling adaptive gradient memory without extra network passes or explicit learning-rate scaling. Experiments on four multimodal benchmarks demonstrate consistent improvements with modest training overhead.