一个适配器,多项任务:面向持续学习的任务条件特征变换
One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual Learning
浏览论文内容
中文总结 AI 辅助
该研究针对持续学习的类增量学习问题,提出FACET方法,通过任务条件特征变换与特征一致性损失,在单适配器下实现高效且抗灾难性遗忘的性能,在长短任务序列上均表现优异。
中文摘要 AI 辅助
类增量学习(CIL)要求模型在不访问早期训练数据的情况下,逐步学习包含新类别的任务,同时保留识别所有已见类别的能力。近期,基于预训练模型的方法通过附加轻量可训练模块适配冻结的骨干网络,已成为主流。然而现有方法存在局限:任务特定适配器学习显式的单任务表示,但参数和计算效率低下;基于LoRA的合并方法将单任务LoRA参数组合为单一模型,其静态聚合权重会在推理时引发表示干扰。为解决这些问题,我们提出FACET:即具有条件特征一致性的任务条件特征变换(Task-conditioned Feature transformation with Conditioned feature consistency),在推理时可生成高判别性特征,同时实现出色的参数效率。在任务序列上持续训练时,FACET学习单个共享适配器,采用动态任务条件特征变换,将适配器的整体特征分布塑造成重叠减少的任务特定组件的混合体。此外,我们提出一种高效的无重放任务条件特征一致性损失,旨在缓解适配器特征空间中学习到的混合分布的灾难性遗忘。即使仅维护单个适配器,FACET也展现出稳健的可扩展性。在极长任务序列(如200个任务)和标准短任务序列(如20个任务)上,我们的方法均实现了优越性能,同时使用的可训练参数和GFLOPs显著更少。代码将在录用后开源。
英文摘要
Class-incremental learning (CIL) requires a model to incrementally learn tasks that contain new classes without accessing earlier training data while preserving the ability to recognize all seen classes. Recently, pretrained-model-based approaches have become prevalent by adapting a frozen backbone with additional lightweight trainable modules. Existing methods, however, exhibit limitations: task-specific adapters learn explicit per-task representations but are parameter- and computation-inefficient, while LoRA-based merging methods combine per-task LoRA parameters into a single model whose static aggregated weights cause representation interference during inference. To address these problems, we present \textbf{FACET}: task-conditioned \textbf{F}e\textbf{A}ture transformation with \textbf{C}ondition\textbf{E}d feature consis\textbf{T}ency, achieving excellent parameter efficiency while producing highly discriminative features during inference. When continually trained on a task sequence, FACET learns a single shared adapter that employs a dynamic task-conditioned feature transformation, shaping the overall feature distribution of the adapter into a mixture of overlap-reduced task-specific components. On the other hand, we propose an efficient replay-free task-conditioned feature consistency loss, aiming to mitigate catastrophic forgetting of the learned mixture distribution in the adapter's feature space. Even when maintaining only a single adapter, FACET demonstrates robust scalability. On both very long task sequences (e.g., 200 tasks) and standard short task sequences (e.g., 20 tasks), our method achieves superior performance while using significantly fewer trainable parameters and GFLOPs. The code will be made open source upon acceptance.
发表机构
- School of Computing and Data Science(计算与数据科学学院)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。