发表机构
Robert Gordon University(罗伯特戈登大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出相对参数重要性度量,用于无重放、无任务ID的持续学习,可平衡稳定性与可塑性并实现反向知识迁移,在两类文本分类任务上优于现有方法
AI 中文摘要
深度神经网络实现持续学习(CL)需要平衡稳定性与可塑性,同时支持知识迁移。本研究聚焦于满足以下约束的离线学习算法:(I)无法访问先前任务的训练数据;(II)推理时无法获取任务ID。我们提出一种新度量——相对参数重要性,该度量衡量每个参数相对于当前及过往任务的重要性。相对重要性高的参数被视为对维持过往任务稳定性更关键,因此会受到严格正则化;而相对重要性低的参数则允许更自由地更新。与现有方法不同,我们的方法允许过往任务重要性高但相对重要性低的参数进行更新,从而在解决稳定性-可塑性权衡问题的同时实现反向知识迁移。我们在类增量和域增量学习文本分类问题上,展示了该方法相对于最先进CL方法的性能提升,并提供了将方法扩展至文本生成问题的见解。代码可在以下URL获取:this https URL
英文摘要
Achieving continual learning (CL) with deep neural networks requires balancing stability and plasticity while enabling knowledge transfer. In this work, we focus on offline learning algorithms under the constraints: (I) no access to training data from prior tasks (II) no access to task-id at inference time. We introduce a novel measure, the relative parameter-importance, which measures the relative importance of each parameter with respect to both the current and past tasks. Parameters with high relative importance are interpreted as more important for maintaining past-task stability and thus heavily regularised, whereas parameters with low relative-importance are allowed to be more freely updated. Unlike existing methods, our approach allows the update of parameters with high past-task importance when they have low relative-importance, thus enabling backward knowledge transfer in addition to tackling the stability-plasticity trade-off. We demonstrate improvements against state-of-the-art CL methods on both class-incremental and domain-incremental learning text classification problems and provide insights for extending our method to text generation problems. Code available at: https://github.com/itsmemala/LACL
CommentsAccepted for publication at the SCL Workshop, ECML-PKDD 2026