发表机构
College of Computer Science, Sichuan University; School of Artificial Intelligence, Sichuan University; Tianfu Jincheng Laboratory; School of Economics and Management, China University of Petroleum (Beijing)(四川大学计算机学院; 四川大学人工智能学院; 天府锦城实验室; 中国石油大学(北京)经济管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模态集合随任务变化的动态多模态持续学习,提出NeuCME框架,结合模态组合重演、多门控专家混合和任务相关性引导蒸馏,在四个真实数据集上显著优于现有方法。
AI 中文摘要
多模态持续学习近来在开发具有类人智能的智能体方面展现出巨大潜力,其通过跨多种模态持续学习新任务来实现。然而,现有方法通常假设每个任务的模态集合是预定义且固定的。在本文中,我们研究了一种更现实的学习设置,称为动态多模态持续学习,其中模态集合可能随任务而变化,而非保持固定。该设置涉及两个主要挑战:(i)时空灾难性遗忘和(ii)自适应多模态融合。为解决这些挑战,我们提出了NeuCME(即多专家神经组合的缩写),这是一种新颖框架,旨在有效学习和整合跨具有不同模态任务的知识。所提出的NeuCME模型包含三个关键组件,即模态组合重演、多门控专家混合和任务相关性引导蒸馏。此外,我们制定了一个评估指标来量化任务序列的动态性,然后建立了一个具有不同动态程度的综合基准。使用四个真实世界数据集的大量实验表明,所提出的NeuCME显著优于最先进的方法。
英文摘要
Multimodal continual learning has recently shown great potential for developing agents with human-like intelligence by continuously learning new tasks across multiple modalities. However, existing methods typically assume that the set of modalities per task is predefined and fixed. In this paper, we investigate a more realistic learning setting, referred to as dynamic multimodal continual learning, in which the set of modalities may vary across tasks rather than remaining fixed. This setting involves two primary challenges: (i) spatio-temporal catastrophic forgetting and (ii) adaptive multimodal fusion. To address these challenges, we propose NeuCME (as shorthand for \textbf{Neu}ral \textbf{C}ombinatorics of \textbf{M}ultiple \textbf{E}xperts), a novel framework designed to effectively learn and integrate knowledge across tasks with varying modalities. The proposed NeuCME model comprises three key components, namely modality-combinational rehearsal, multi-gated mixture-of-experts, and task relevance-guided distillation. Furthermore, we formulate an evaluation metric to quantify the dynamism of task sequences and then set up a comprehensive benchmark with different degrees of dynamism. Extensive experiments using four real-world datasets demonstrate that the proposed NeuCME outperforms state-of-the-art methods markedly.
Comments9 pages, 5 figures