arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过教师预测精化增强检测Transformer的知识蒸馏

Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement

Yitong Xing, Yuhao Cheng, Yanping Li, Yichao Yan

arXiv 2609.19964首次发表:更新:

发表机构

Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对DETR蒸馏中教师监督质量不佳的问题,提出TPRD模块,通过正预测校正和负预测抑制精化教师预测,并引入MDKP保留暗知识,在MS COCO和PASCAL VOC上验证了有效性。

AI 中文摘要

检测Transformer(DETR)在目标检测中取得了强劲性能,但由于其高计算成本,在边缘设备上部署仍具挑战性。现有的DETR蒸馏方法主要关注对齐蒸馏点,而很大程度上忽视了教师监督本身的质量。我们观察到,由于DETR中阶段性的非单调预测行为,早期阶段定位良好或分类正确的预测可能在后期阶段退化,且一些负预测变得日益过度自信。因此,仅依赖当前阶段的预测会产生不准确且不一致的监督。为解决此问题,我们提出了教师预测精化蒸馏(TPRD),这是一种即插即用模块,通过利用阶段性预测信息在蒸馏前精化教师预测。TPRD通过正预测校正(PPC)提高监督质量,该校正通过恢复早期阶段更准确的预测来修正退化的正预测,确保可靠的定位和分类信号;负预测抑制(NPS)则抑制过度自信负预测的影响,防止它们向学生提供误导性监督。为保留有信息的暗知识,我们进一步引入了最大暗知识保留(MDKP),该机制选择性地精化目标类别的logits,同时保留非目标类别间的关系。在MS COCO和PASCAL VOC上的大量实验证明了所提方法的有效性和鲁棒性。我们的代码可在https://this URL获取。

英文摘要

Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher's supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage's predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at https://github.com/xingyitong1/TPRD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑