arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20302cs.LGcs.MM

SAGG:样本自适应梯度门控用于异构损坏下的鲁棒多模态学习

SAGG: Sample-Adaptive Gradient Gating for Robust Multimodal Learning under Heterogeneous Corruption

  • Tsinghua University(清华大学)
  • Sichuan University(四川大学)

机构由 AI 辅助整理,请以论文原文为准。

Wentao Zhang, Yifan Zhu, Yutong Zhang, Wentao Mo

AI总结:

针对多模态学习中样本级异构损坏问题,提出样本自适应梯度门控(SAGG),通过逐样本二值门控与截断机制实现无偏梯度估计,理论保证收敛性,实验在多个数据集上优于现有方法。

AI中文摘要:

多模态梯度平衡方法通过每个模态的共享标量来调节编码器梯度,隐含地假设训练批次中的损坏是均匀的。实际上,损坏是样本异构的:在单个小批次内,不同样本可能具有不同的模态被损坏。我们证明,在这种异构损坏模型下,任何具有共享调制参数的批次级样本无关线性估计器相对于干净数据梯度都会产生不可约的偏差,并且在自然无分布估计器类中,样本级全有或全无门控是唯一无偏策略。受此结果启发,我们提出了样本自适应梯度门控(SAGG),它通过在线特征范数质量测试对每个样本做出二进制的保留或丢弃决策,并结合截断机制进行方差控制。我们证明,基于SAGG的随机梯度下降以标准的O(1/sqrt(T))速率收敛到干净损失的驻点,且没有依赖于损坏的错误底线,并为独立编码器架构推导了认证鲁棒半径,该半径将每个模态的Lipschitz常数与分类间隔联系起来。在Kinetics-Sounds和UCF-101上,在高斯噪声注入、部分模态缺失和自然贡献不平衡下进行的实验表明,SAGG始终优于十种现有方法,在批次级偏差最严重的高损坏情况下增益最大。

英文摘要:

Multimodal gradient balancing methods modulate encoder gradients with a shared scalar per modality, implicitly assuming that corruption is uniform across the training batch. In practice, corruption is sample-heterogeneous: within a single mini-batch, different samples may have different modalities corrupted. We prove that under this heterogeneous corruption model, any batch-level sample-agnostic linear estimator with a shared modulation parameter incurs an irreducible bias with respect to the clean-data gradient, and that sample-level all-or-nothing gating is the unique unbiased strategy within a natural distribution-free estimator class. Motivated by this result, we propose Sample-Adaptive Gradient Gating (SAGG), which makes a binary retain-or-discard decision per sample via an online feature-norm quality test and incorporates a truncation mechanism for variance control. We prove that SAGG-based SGD converges at the standard O(1/sqrt(T)) rate to stationary points of the clean loss without a corruption-dependent error floor, and derive a certified robustness radius for the independent-encoder architecture that connects per-modality Lipschitz constants to the classification margin. Experiments on Kinetics-Sounds and UCF-101 under Gaussian noise injection, partial modality missing, and natural contribution imbalance show that SAGG consistently outperforms ten existing methods, with the largest gains in high-corruption regimes where batch-level bias is most severe.

补充信息

↑