arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多模态无监督持续后训练的视觉依赖感知框架

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu

arXiv 2608.26095首次发表:更新:

发表机构

State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences; Pengcheng Laboratory; School of Artificial Intelligence, University of Chinese Academy of Sciences; Harbin Institute of Technology, Shenzhen(中国科学院自动化研究所多模态人工智能系统国家重点实验室; 鹏城实验室; 中国科学院大学人工智能学院; 哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态无监督持续后训练中现有方法忽略视觉依赖的问题,提出含VC-OT和VMA组件的VDA框架,可同时维持旧任务稳定性与新任务可塑性,经实验验证有效。

AI 中文摘要

本文探索多模态无监督持续后训练(Multimodal Unsupervised Continual Post-Training, MU-CPT)这一新任务,使已部署的多模态大语言模型(Multimodal Large Language Models, MLLMs)能从流式无标签数据中持续进化。现有针对MLLMs的无监督后训练方法通常对目标词元进行均匀优化,忽略了它们的异构视觉依赖(Visual Dependence, VD)。然而我们发现,词元级VD对MU-CPT至关重要:其结构扭曲是跨模态灾难性遗忘的指标,而其内在异质性则是引导新任务学习的方向。利用这一特性,我们提出视觉依赖感知(Visual Dependence-Aware, VDA)框架,包含两个核心组件:其一,视觉约束最优传输(Visually Constrained Optimal Transport, VC-OT)将新任务学习过程中旧任务VD的结构扭曲建模为最优传输问题,以缓解跨模态遗忘;通过设计区域感知的基础代价和依赖分层的传输惩罚,它可防止视觉焦点的全局偏移,同时严格禁止视觉依赖退化为语言偏差。其二,视觉调制适配(Visually Modulated Adaptation, VMA)利用VD异质性来强化基于视觉的新任务学习,促进新任务可塑性。我们的方法在具有挑战性的MU-CPT场景下,同时维持了旧任务稳定性和新任务可塑性。在我们的MU-CPT设置下开展的大量实验验证了VDA的有效性。

英文摘要

In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting, and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑