arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18716cs.CV

在不断演变的领域上进行连续视频 - MLLM 适应

Continual Video-MLLM Adaptation over Evolving Domains

Rui Cheng, Meixing Shi, Yuxiang Cai, Jingcai Guo, Jianwei Yin, Zhi Chen

首次发表
浏览论文内容

中文总结 AI 辅助

研究视频多模态大语言模型对不断演变领域的适应问题,提出分布感知专家路由框架 DAER,通过保持领域隔离的轻量级专家、引入多种路由和优化机制,在域增量基准上评估,结果优于现有方法。

中文摘要 AI 辅助

视频多模态大语言模型在视频理解方面展现出强大能力,但对不断演变的领域的适应仍未得到充分探索。在实际部署中,视频数据常从异构领域持续到达,要求模型获取新的领域特定知识而不覆盖先前学到的能力。现有连续学习方法依赖共享适应空间,会引发严重跨域干扰和灾难性遗忘。我们提出分布感知专家路由(DAER),这是一个用于在不断演变的领域上进行连续视频 - MLLM 适应的参数高效框架。DAER 保持领域隔离的轻量级专家,同时冻结预训练的视频 - MLLM 主干,将特定领域适应与预训练模型的通用多模态知识解耦。为实现细粒度专业化,引入域内分布感知路由机制,通过最大均值差异(MMD)将每个输入与专家级原型库匹配。为解决推理时任务标识缺失问题,进一步提出域间路由机制,在判别性子空间中进行原型匹配以实现鲁棒的领域识别。此外,引入自适应域合并以提高参数可扩展性,并采用两阶段优化策略在连续学习期间稳定专家专业化。我们通过构建一个由十个涵盖不同视觉环境和推理需求的 VidQA 数据集组成的域增量基准来评估 DAER。在两个强大的视频 - MLLM 主干上的实验表明,DAER 始终优于先前方法。

英文摘要

Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remains underexplored. In real-world deployments, video data often arrives continuously from heterogeneous domains, requiring the model to acquire new domain-specific knowledge without overwriting previously learned capabilities. Existing continual learning methods typically rely on shared adaptation spaces, which can induce severe cross-domain interference and catastrophic forgetting. We propose Distribution-Aware Expert Routing, a parameter-efficient framework for continual Video-MLLM adaptation over evolving domains. DAER maintains domain-isolated lightweight experts while keeping the pretrained Video-MLLM backbone frozen, thereby decoupling domain-specific adaptation from the general multimodal knowledge of the pretrained model. To enable fine-grained specialization, we introduce an intra-domain distribution-aware routing mechanism that matches each input to expert-level prototype reservoirs using MMD. To address the absence of task identities at inference time, we further propose an inter-domain routing mechanism that performs prototype matching in a discriminative subspace for robust domain identification. In addition, we introduce adaptive domain merging to improve parameter scalability and adopt a two-stage optimization strategy to stabilize expert specialization during continual learning. We evaluate DAER by curating a domain-incremental benchmark built from ten VidQA datasets covering diverse visual environments and reasoning demands. Experiments on two strong Video-MLLM backbones show that DAER consistently outperforms prior methods.

发表机构

  • School of Software Technology, Zhejiang University(浙江大学软件学院)
  • Hong Kong Polytechnic University(香港理工大学)
  • The University of Southern Queensland(南昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑