CARE:基于师生知识蒸馏的细粒度视觉分类的约束注意力细化
CARE: Constrained Attention Refinement for Fine-Grained Visual Classification via Teacher-Student Distillation
- College of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对细粒度视觉分类中注意力通路判别能力不足的问题,提出CARE框架,通过师生知识蒸馏结合正则化项提升性能与可解释性,在多个数据集上取得良好效果。
AI中文摘要:
细粒度视觉分类要求模型识别细微的局部特征,同时揭示其预测背后的视觉证据。特定类别的注意力通路为可解释识别提供了自然基础,但其受约束的预测结构限制了判别能力,且未充分利用强大预训练骨干网络的中间表示。为解决该问题,我们提出CARE,这是一种通过师生知识蒸馏实现可解释细粒度识别的约束注意力细化框架。CARE将最终预测和解释保留在特定类别的注意力学生模型中,同时引入仅用于训练的辅助查询教师模型,该模型使用可学习查询读取选定的中间DINOv2层。教师模型融合多级表示,并将经过logit标准化的类别判别知识传递给学生模型。为进一步细化解释通路,我们设计了多样性和稀疏性项来正则化学生注意力头,减少冗余并鼓励紧凑的特征定位。在CUB、Oxford-IIIT Pet、Stanford Dogs和Stanford Cars数据集上的实验表明,CARE在可解释的冻结骨干网络设置下实现了强劲的分类性能,在CUB上达到78.5%的Top-1准确率。使用插入和删除指标进行的忠实性分析进一步表明,排名靠前的注意力区域保留了与类别相关的解释证据。
英文摘要:
Fine-grained visual classification requires models to recognize subtle local traits while exposing the visual evidence behind their predictions. Class-specific attention pathways provide a natural basis for interpretable recognition, but their constrained prediction structure limits discriminative capacity and underuses intermediate representations from strong pretrained backbones. To address this problem, we propose CARE, a constrained attention refinement framework for interpretable fine-grained recognition via teacher-student distillation. CARE keeps the final prediction and explanation within a class-specific attention student, while introducing a training-only auxiliary query teacher that reads selected intermediate DINOv2 layers with learnable queries. The teacher fuses multi-level representations and transfers logit-standardized class-discriminative knowledge to the student. To further refine the explanation pathway, we design diversity and sparsity terms to regularize student attention heads, reducing redundancy and encouraging compact trait localization. Experiments on CUB, Oxford-IIIT Pet, Stanford Dogs, and Stanford Cars show that CARE achieves strong classification performance under an interpretable frozen-backbone setting, reaching 78.5% Top-1 accuracy on CUB. Faithfulness analysis with insertion and deletion metrics further indicates that the top-ranked attention regions retain class-relevant evidence for explanation.