arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36759cs.CVcs.AI

双模式低秩学习器与桥接原型集成用于视觉-语言类增量学习

Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

Chiyuan He, Zihuan Qiu, Fanman Meng, Chao Wang, Liangjiang Chen, Linfeng Xu, Qingbo Wu, Hongliang Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对CLIP类增量学习中知识覆盖与模态整合不足的问题,提出DuLBE,通过双模式低秩学习和桥接原型集成,实现无样本CIL,在保持参数效率的同时达到最先进性能。

中文摘要 AI 辅助

得益于可迁移的视觉-文本对齐,CLIP已被广泛用于类增量学习(CIL)。然而,现有学习器要么反复更新跨任务共享的组件,导致知识覆盖,要么过度隔离新任务更新,阻碍了CLIP可迁移知识的复用并限制了可塑性。此外,基于文本或双模态的分类器设计仍未能有效整合视觉和文本模态的互补信息。为应对这些挑战,我们提出了DuLBE,它将双模式低秩学习与桥接原型集成分类器相结合,用于无样本CIL。DuLBE根据梯度需求分配两种视觉低秩更新模式,并使用梯度路由进行协调:从历史占用的视觉方向中选择一个紧凑且可重写的共享模式以复用可迁移知识,而残差模式为任务特定变化提供低干扰通道。基于由此产生的稳定跨模态结构,我们进一步在单位超球面上构建视觉原型与文本嵌入之间的测地线桥,并集成可靠的桥接原型以补偿文本决策边界的模态差距限制。在多种设置下的大量实验表明,DuLBE在保持低秩调优高参数效率的同时,实现了最先进的CIL性能。

英文摘要

Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable knowledge and limiting plasticity. Moreover, the text-based or bimodal classifier designs still fail to effectively integrate complementary information from the visual and textual modalities. To address these challenges, we introduce DuLBE, which couples dual-mode low-rank learning with a bridge-prototype ensemble classifier for exemplar-free CIL. DuLBE allocates two visual low-rank update modes according to the gradient demand and uses gradient routing to coordinate them: a compact and rewritable shared mode is selected from historically occupied visual directions to reuse transferable knowledge, while residual modes provide low-interference channels for task-specific variations. Building on the resulting stable inter-modal structure, we further construct geodesic bridges between visual prototypes and text embeddings on the unit hypersphere, and ensemble reliable bridge prototypes to compensate for the modality-gap limitations of textual decision boundaries. Extensive experiments under multiple settings show that DuLBE achieves state-of-the-art CIL performance while retaining the high parameter efficiency of low-rank tuning.

发表机构

  • University of Electronic Science and Technology of China(电子科技大学)
  • Qiyuan Lab(启元实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑