发表机构
Beijing University of Posts and Telecommunications; Beihang University; Beijing University of Technology(北京邮电大学; 北京航空航天大学; 北京工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出诊断框架,区分可校正偏差与类别对应关系,发现偏移可校正删除后来类别匹配的大代价,而破坏类别对应关系的小代价持续存在。
AI 中文摘要
类增量学习必须在没有任务标签的情况下识别迄今为止见过的所有类别。诸如DER和DER++之类的Logit重放方法通过匹配模型在存储示例上的过去预测来减轻遗忘。删除这种匹配揭示了其益处,但由此产生的准确率损失无法表明存储的分数本身是否必要,或者该损失是否能在校正分类器对近期类别的偏差后得以消除。我们提出一个诊断框架,将缓存的预测视为时间上异构的监督:它区分了示例存储时已知的类别与之后学习的类别,对每组进行编辑,并在任务级偏移前后评估每个模型,该偏移保持任务内预测不变。在CIFAR-100上使用DER++,合适的固定常数可在1个百分点的等价裕度内替换后来学习类别的未刷新存储分数,并且该偏移将删除其匹配的代价从14.9点降至1.8点。重新分配存储时已知类别的非金标分数(保留其值和每个任务的目标概率)在偏移前代价为4.3点,偏移后为4.0点,并且图像蒸馏中存在平行的代价。在测试的固定头部设置中,删除后来类别匹配的大代价因此大部分可由该偏移校正,而破坏类别对应关系的较小代价则持续存在。代码和数据可在该HTTP URL获取。
英文摘要
Class-incremental learning must recognize all classes seen so far without task labels. Logit replay methods such as DER and DER++ mitigate forgetting by matching the model's past predictions on stored examples. Deleting this matching reveals its benefit, but the resulting accuracy cost cannot show whether the stored scores themselves are needed, or whether the cost survives correction of the classifier's bias toward recent classes. We propose a diagnostic framework that treats a cached prediction as temporally heterogeneous supervision: it separates classes known when an example was stored from classes learned afterward, edits each group, and evaluates every model before and after a task-level offset that leaves within-task predictions unchanged. On CIFAR-100 with DER++, suitable fixed constants replace the unrefreshed stored scores of later-learned classes within an equivalence margin of 1 percentage point, and the offset reduces the cost of deleting their matching from 14.9 to 1.8 points. Reassigning the non-gold scores of classes known at storage, which preserves their values and each task's target probability, costs 4.3 points before and 4.0 after the offset, and a parallel cost persists in image distillation. In the tested fixed-head setting, the large cost of deleting later-class matching is thus mostly correctable by this offset, whereas the smaller cost of disrupting class correspondence persists. Code and data are available at anonymous.4open.science/r/replay-preserve-E22B.