arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TASSO:面向视觉-语言模型持续学习的任务特定子空间优化

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

Chang Sun, Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh

arXiv 2608.21487首次发表:更新:

发表机构

University of Padova(帕多瓦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

TASSO通过子空间学习与几何感知知识蒸馏,缓解VLMs持续学习中的灾难性遗忘和零样本退化,在多域及类增量学习基准上优于现有最优方法。

AI 中文摘要

视觉-语言模型(Vision-Language Models, VLMs)具备强大的零样本能力,是跨多样任务持续学习的理想方案。但在持续适配过程中,会同时出现灾难性遗忘和零样本性能退化问题,严重降低表现。本文提出TASSO这一新范式,可在高效保留隐空间几何结构的同时确保网络可塑性,通过两种互补技术实现:子空间学习与几何感知知识蒸馏。具体而言,我们首先学习一系列任务特定的低秩投影器,用于在优化交叉熵前对隐表征进行投影;其次,采用基于测地距离的损失函数,从之前任务的模型中蒸馏知识,同时有效保留隐空间几何结构。这些设计选择不仅避免了在完整嵌入维度上进行不必要的参数更新,还通过聚焦任务特定流形提升了学习效果;此外,几何感知蒸馏提供了强正则化,显著降低了持续学习过程中的灾难性遗忘和零样本退化。在CLIP视觉-语言模型上开展的多域任务增量学习与类增量学习基准实验结果显示,该方法在缓解遗忘和保留零样本能力方面,较现有最优方法取得了明显提升。

英文摘要

Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diverse tasks. However, during continual adaptation, both catastrophic forgetting and zero-shot degradation occur, severely degrading performance. In this paper, we introduce TASSO, a new paradigm that efficiently preserves the latent space geometry while ensuring network plasticity. We achieve this with two complementary techniques: subspace learning and geometry-aware knowledge distillation. Specifically, we first learn a sequence of task-specific low-rank projectors, which we use to project the latent representations before optimizing cross-entropy. Secondly, we employ a geodesic-distance-based loss that distills knowledge from the previous-task model while effectively preserving the latent space geometry. These design choices not only avoid unnecessary parameter updates along the full embedding dimensions but also improve learning by focusing on task-specific manifolds. Moreover, the geometry-aware distillation provides strong regularization and significantly reduces both catastrophic forgetting and zero-shot degradation throughout the continual learning sequence. Experimental results with the CLIP vision language model in the multi-domain task incremental and class incremental learning benchmarks demonstrate clear improvements over state-of-the-art methods in mitigating forgetting and preserving zero-shot capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑