基于视觉基础模型的跨域目标检测的语义兼容知识蒸馏
Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models
- National Supercomputing Center in Changsha, Hunan University(湖南大学长沙国家超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对跨域目标检测中VFM的语义不兼容问题,提出SLE-T框架,通过适配DINOv2实现高效知识迁移,在DAOD基准上达到SOTA性能,且训练成本更低。
AI中文摘要:
视觉基础模型(Vision Foundation Models, VFMs)为域自适应目标检测(Domain-Adaptive Object Detection, DAOD)提供了强大的泛化能力。然而,现有基于VFM的方法忽略了教师与学生特征图之间的空间尺度差异,导致语义不兼容,削弱了特征对齐和伪标签学习。此外,域偏移会导致源域训练的VFM教师遗漏目标域对象,限制了其伪标签的质量。为解决这些问题,我们提出了语义定位增强教师(Semantic Localization-Enhanced Teacher, SLE-T),这是一个围绕轻量SLE Adapter构建的语义兼容知识蒸馏框架,适配于DINOv2。SLE Adapter将预训练的局部纹理先验注入DINOv2,以提升跨域识别能力,并将其特征重构为与学生检测器在空间和语义上兼容的密集表示。SLE-T通过伪标签学习或特征对齐传递得到的教师知识。我们用DINOv2-B和DINOv2-L(ViT-B和ViT-L变体)实例化SLE-T,并与更大的DINOv2-G教师进行比较。在三个DAOD基准上的大量实验表明,我们的方法达到了最先进的性能, ablation研究证实了教师-学生语义兼容性的重要性。值得注意的是,使用DINOv2-B的SLE-T生成的伪标签具有竞争力或更优,且训练时间约为DINOv2-G的四分之一,GPU内存消耗也显著更少,证明了在有限计算资源下的高效VFM知识迁移。
英文摘要:
Vision foundation models (VFMs) offer strong generalization capabilities for domain-adaptive object detection (DAOD). However, existing VFM-based methods overlook the spatial-scale discrepancy between teacher and student feature maps, resulting in semantic incompatibility that weakens both feature alignment and pseudo-label learning. Moreover, domain shift can cause source-trained VFM teachers to miss target-domain objects, limiting the quality of their pseudo-labels. To address these issues, we propose the Semantic Localization-Enhanced Teacher (SLE-T), a semantically compatible knowledge-distillation framework built around a lightweight SLE Adapter for DINOv2. SLE Adapter injects pretrained local-texture priors into DINOv2 to improve cross-domain recognition and reformulates its features into dense representations that are spatially and semantically compatible with the student detector. SLE-T transfers the resulting teacher knowledge through either pseudo-label learning or feature alignment. We instantiate SLE-T with DINOv2-B and DINOv2-L (the ViT-B and ViT-L variants) and compare them with the larger DINOv2-G teacher. Extensive experiments on three DAOD benchmarks demonstrate that our method achieves state-of-the-art performance, and ablation studies confirm the importance of teacher-student semantic compatibility. Notably, SLE-T with DINOv2-B produces competitive or superior pseudo-labels using approximately one-quarter of the training time of DINOv2-G and substantially less GPU memory, demonstrating efficient VFM knowledge transfer under limited computational resources.