arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态持续学习中模态贡献漂移的正则化

Regularizing modality contribution drift in multimodal continual learning

Zhen Zhang, Jielei Chu, Wenjie Ban, Tian Sang, Yuxiao Li, Bin Liu, Fengmao Lv, Tianrui Li

arXiv 2607.27260首次发表:更新:

发表机构

School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态持续学习中的模态贡献漂移问题,提出含基于重放和无重放版本的CMCDR方法,经实验验证其通用性与有效性。

AI 中文摘要

多模态持续学习(MMCL)旨在从多模态数据中学习新知识,同时保留已有知识。为缓解遗忘,现有MMCL方法通常关注跨模态表示对齐或语义相似性,但忽略了各模态的相对贡献及其交互在增量任务中是否保持稳定,我们将这种决策层面的偏移称为模态贡献漂移(MCD),并通过MCD分数对其进行量化,该分数结合了在模态子集受控干预下的贡献强度和相对依赖变化。理论与实证分析进一步解释了为何现有MMCL方法无法可靠缓解这种漂移。为此,我们提出持续模态贡献漂移正则化(CMCDR),其保留了先前学习任务的模态贡献结构。由于MMCL设置在是否有旧样本可用方面存在差异,CMCDR包含基于重放和无重放两种版本:基于重放的版本使用模态子集干预作为对存储旧样本的诊断探针,比较当前模型与冻结旧模型间的贡献分布,并约束旧样本的模态特定及交互贡献的变化;无重放版本使用当前任务样本作为探针,蒸馏冻结模型的旧任务贡献响应,从而在无样本时对观测到的贡献分布进行正则化。在多模态类增量学习和持续视觉问答上的实验验证了CMCDR的通用性与有效性。

英文摘要

Multimodal continual learning (MMCL) aims to acquire new knowledge from multimodal data while retaining previously learned knowledge. Existing MMCL methods primarily mitigate forgetting by aligning cross-modal representations or preserving feature-level semantic similarity. However, different tasks may rely on different modalities, and learning new tasks can alter how modalities contribute to predictions on previously learned tasks. It remains underexplored how modality contributions evolve across incremental stages and how such changes relate to forgetting in MMCL. We term such changes Modality Contribution Drift (MCD) and introduce an MCD score based on controlled modality-subset interventions. Our theoretical and empirical analyses show how contribution drift can lead to forgetting, while existing MMCL and conventional CL methods do not effectively mitigate MCD. To address this issue, we propose Continual Modality Contribution Drift Regularization (CMCDR) to preserve the modality contribution profiles of previously learned tasks in both replay-based and replay-free settings. In the replay-based setting, CMCDR estimates modality contributions from stored old samples and regularizes their drift relative to a frozen previous model. In the replay-free setting, CMCDR uses current-task samples as probes to match old-class contribution profiles between the current and frozen models without storing old exemplars. Across six benchmarks covering multimodal class-incremental learning and continual multimodal question answering, CMCDR substantially reduces modality contribution drift, improves average accuracy by 0.88-7.66 percentage points, and reduces average forgetting by 0.85-11.55 percentage points over the corresponding baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑