TOOD:面向持续学习智能体的任务感知分布外分数校准
TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners
浏览论文内容
中文总结 AI 辅助
本文针对持续学习智能体的分布外检测遗忘问题,提出无需训练的TOOD事后校准方法,在多数据集上显著提升了分布外检测性能。
中文摘要 AI 辅助
持续学习(CL)系统的核心挑战是在学习新任务的同时保持对已学任务的性能。持续学习系统另一个同样重要但研究较少的方面是区分输入是否来自系统已遇任务集合之外的能力,即分布外(OOD)检测。本文报告了持续学习系统中OOD检测动态、随时间性能下降原因(称为OOD遗忘,OODF)及缓解策略的多项发现。主要发现:OODF与旧任务分类性能仅呈弱负相关,说明二者机制不同,且该效应对基于能量和基于特征的OOD检测方法均存在。基于能量的检测器因学习更多任务出现logit尺度下降,称为置信间隙;基于特征的检测器则因互补效应(称为流形拥挤)性能下降。基于上述观察,提出TOOD,一种无需训练的事后方法,将logits分解为单任务能量分数,并用回放缓冲区统计数据重新校准。在CIFAR-10、CIFAR-100及100任务ImageNet-1K流上的实验显示,多数设置下TOOD的OOD检测性能优于未校准的能量方法,在10个CIFAR配置中9个排第一或第二,置信间隙最严重时增益最大。结果表明,持续学习中OOD恶化的很大一部分源于分数校准不当,而非判别结构完全丢失。
英文摘要
The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly important though less well-studied aspect of CL systems is their ability to distinguish inputs that are unlikely to come from within the set of tasks the system has already encountered, often called out-of-distribution (OOD) detection. This paper presents several findings related to the dynamics of OOD detection in CL systems, causes of performance degradation over time which we call OOD forgetting (OODF), and proposed mitigation strategies for this degradation. Chiefly, we find the unintuitive result that OODF is only weakly anti-correlated with classification performance on previous tasks, suggesting that the underlying mechanisms producing OODF are distinct. Moreover, this effect is observed for both energy-based and feature-based OOD detection methods. Energy-based detectors suffer a drop in logit scale as additional tasks are learned, which we term the Confidence Gap, while feature-based detectors also degrade under a complementary effect we call Manifold Crowding. Motivated by these observations, we propose TOOD, a training-free post-hoc method that decomposes logits into per-task energy scores and re-calibrates them using replay-buffer statistics. Experiments on CIFAR-10, CIFAR-100, and a 100-task ImageNet-1K stream show that TOOD improves OOD detection performance over uncalibrated energy in most settings and ranks first or second in nine of ten CIFAR configurations, with the largest gains when the confidence gap is most severe. These results suggest that a substantial portion of OOD deterioration in continual learning arises from score miscalibration rather than from a complete loss of discriminative structure.
发表机构
- Mila - Quebec AI Institute(米拉-魁北克人工智能研究所)
- CIFAR AI Chair(加拿大高级研究所人工智能主席项目)
机构由 AI 辅助整理,请以论文原文为准。