arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37888cs.CVcs.LG

视觉分支是CLIP类增量学习所需的关键

Visual Branch is What You Need for CLIP-based Class-Incremental Learning

Tao Hu, Zhen-Hao Xie, Jingcai Guo, De-Chuan Zhan, Da-Wei zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对CLIP类增量学习中文本分支因模态差距而性能不佳的问题,提出纯视觉方法VIS,利用视觉特征构建增量分类器,实现高效更新并达到最优性能。

中文摘要 AI 辅助

类增量学习(CIL)要求模型在不遗忘先前所学类别的情况下,随时间识别新类别。随着视觉-语言预训练的兴起,CLIP已成为CIL的强有力基础。基于CLIP的CIL中一种常见设计是,通过使用CLIP文本编码器对类名模板进行编码来构建文本分类器权重,然后通过图像-文本余弦相似度对视觉特征进行分类。这种设计颇具吸引力:由于CLIP在共享嵌入空间中对齐图像和文本,文本权重似乎为增量类别提供了现成的分类器。然而,我们表明这种看似自然的设计并非总是有益的,因为模态差距仍可能将两种模态分开,使文本分类器权重偏离视觉类别分布。实验上,在相同的任务级CIL训练下,使用视觉类别中心初始化余弦分类器比使用CLIP文本权重产生更低的损失和更好的增量准确率。基于这些观察,我们提出VIS,一种用于基于CLIP的CIL的纯视觉方法,它移除部署的文本分支,并完全在视觉空间中构建增量分类器。为了获得更强的任务自适应视觉表示,VIS仅使用基础会话数据,通过信息丰富的视觉层特征增强CLIP的最终视觉表示。基于增强的视觉表示,VIS采用简单的核化增量最小二乘支持向量机,其分类器权重从加性充分统计量中闭式求解。当新类别到达时,VIS累积其充分统计量并重新计算所有已见类别的分类器权重,实现高效的增量更新,同时保留历史类别知识。大量实验表明,VIS在没有文本分支的情况下实现了最先进的性能。

英文摘要

Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textual classifier weights deviate from visual class distributions. Empirically, under identical task-wise CIL training, initializing the cosine classifier with visual class centers yields lower loss and better incremental accuracy than using CLIP textual features. Motivated by these observations, we propose VIS, a visual-only method for CLIP-based CIL that removes the deployed textual branch and constructs the incremental classifier entirely in the visual space. To obtain stronger task-adaptive visual representations, VIS uses only base-session data to enhance CLIP's final visual representation with informative visual-layer features. Built on the enhanced visual representation, VIS employs a simple kernelized incremental least-squares SVM, whose classifier weights are solved in closed form from additive sufficient statistics. When new classes arrive, VIS accumulates their sufficient statistics and recomputes the classifier weights for all seen classes, enabling efficient incremental updates while preserving historical class knowledge. Extensive experiments show that VIS achieves state-of-the-art performance without a textual branch.

发表机构

  • Nanjing University(南京大学)
  • Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑