发表机构
Meituan Vision AI Department(美团视觉AI部门)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对类增量学习中共享表示退化与任务间干扰问题,提出解耦表示动态网络(DRDN),通过掩码图像建模保持骨干网络通用视觉结构,并采用层次化任务令牌扩展减少跨任务干扰,在从头训练的ViT上取得显著提升。
AI 中文摘要
类增量学习(CIL)的动态扩展方法通过增加专用令牌或子网络来保护任务特定知识,但我们的分析表明,仅靠分类监督不足以在长增量序列中充分保留任务无关的共享骨干表示。我们识别出两个相互交织的挑战:一是由于在主要包含当前任务数据上的顺序训练导致的跨任务混淆,这使决策边界偏向近期任务;二是骨干网络中欠优化的共享表示限制了随着任务积累的长期判别能力。我们提出了解耦表示动态网络(DRDN),通过两种正交机制应对这些挑战。对于共享骨干表示,DRDN在每个增量步骤中持续应用掩码图像建模(MIM),重建梯度仅通过骨干网络传播,鼓励其保留超越类别判别线索的通用视觉结构。对于任务特定判别,DRDN在所有Transformer层中采用层次化任务令牌扩展,并采用修改后的每任务注意力规则以减少任务间干扰。我们通过准确率退化分析和跨任务混淆率测量支持了这一设计。在从头训练的ViT CIL设置(无外部预训练)中,DRDN在相同骨干规模下持续优于强令牌扩展基线。在CIFAR100-B0(10步)上,DRDN达到77.19%的平均准确率,比DKT高出1.36个百分点,比DyTox高出3.53个百分点,且优势在更长增量序列中扩大。多种子验证确认了稳定性(+/-0.31%)。MIM解码器仅在训练期间激活,不增加推理时的参数或计算量。
英文摘要
Dynamic expansion methods for class-incremental learning (CIL) protect task-specific knowledge by growing dedicated tokens or subnetworks, yet our analyses suggest that classification supervision alone does not sufficiently preserve task-agnostic shared backbone representations over long incremental sequences. We identify two intertwined challenges: cross-task confusion from sequential training on predominantly current-task data, which biases decision boundaries toward recent tasks; and under-optimized shared representations in the backbone that cap long-term discriminability as tasks accumulate. We propose the Decoupled Representation Dynamic Network (DRDN), which addresses these challenges via two orthogonal mechanisms. For shared backbone representations, DRDN continuously applies masked image modeling (MIM) at every incremental step, with reconstruction gradients routed exclusively through the backbone, encouraging it to retain general visual structure beyond class-discriminative cues. For task-specific discrimination, DRDN employs hierarchical task token expansion across all transformer layers, with a modified per-task attention rule that reduces inter-task interference. We support this design with accuracy degradation analysis and cross-task confusion rate measurements. In the from-scratch ViT CIL setting (no external pretraining), DRDN consistently improves over strong token-expansion baselines with comparable backbone scale. On CIFAR100-B0 (10 steps), DRDN achieves 77.19% average accuracy, outperforming DKT by 1.36 points and DyTox by 3.53 points, with an advantage that grows at longer incremental sequences. Multi-seed validation confirms stability (+/-0.31%). The MIM decoder is active only during training, adding no inference-time parameters or computation.
Comments10 pages, IEEEtran journal format. Preprint submitted to IEEE Transactions on Multimedia