发表机构
National Engineering Lab for Big Data Analytics, School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院大数据分析国家工程实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对持续学习中提示微调的检索依赖和分类器偏差问题,提出解耦提示微调方法,将提示解耦为共享分布与类别特定提示,并给出信息论理论保障,在标准基准上达到最先进性能。
AI 中文摘要
持续学习(CL)旨在从序列数据中增量获取知识,同时避免灾难性遗忘。近年来,提示微调作为一种将预训练模型适配到持续学习任务的高效方法,受到了越来越多的关注。然而,现有的提示设计范式普遍存在检索依赖和分类器偏差问题,这使得模型适配对提示选择敏感,并使预测偏向于新到达的类别。为应对这些挑战,我们提出了用于持续学习的解耦提示微调(DPT4CL),该方法将CLIP文本提示解耦为任务共享提示分布和类别特定提示。任务共享提示分布通过优化信息瓶颈目标来推导,以促进跨任务知识迁移并缓解分类器偏差,而类别特定提示则在不依赖显式提示检索的情况下增强类间可分性。此外,我们从信息论角度建立了统一的超额风险界,为所提出框架的鲁棒泛化和遗忘缓解提供了理论支持。在标准持续学习基准上的大量实验表明,DPT4CL达到了最先进的性能。源代码可在以下网址获取:此HTTPS URL。
英文摘要
Continual learning (CL) aims to incrementally acquire knowledge from sequential data while avoiding catastrophic forgetting. Recently, prompt tuning has attracted increasing attention as an efficient approach for adapting pre-trained models to CL tasks. However, existing prompt design paradigms commonly suffer from retrieval dependence and classifier bias, which make model adaptation sensitive to prompt selection and bias predictions toward newly arrived classes. To address these challenges, we propose Decoupled Prompt Tuning for Continual Learning (DPT4CL), which decouples the CLIP textual prompt into a task-shared prompt distribution and class-specific prompts. The task-shared prompt distribution is derived by optimizing an Information Bottleneck objective to facilitate cross-task knowledge transfer and alleviate classifier bias, while class-specific prompts enhance inter-class separability without relying on explicit prompt retrieval. Furthermore, we establish a unified excess risk bound from an information-theoretic perspective, providing theoretical support for the robust generalization and forgetting mitigation of the proposed framework. Extensive experiments on standard CL benchmarks demonstrate that DPT4CL achieves state-of-the-art performance. The source code is available at https://github.com/Cloudfly-Z/DPT4CL