arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FD-CanKD:用于紧凑目标检测器的精细化先验的频率解耦交叉注意力知识蒸馏

FD-CanKD: Frequency-Decoupled Cross-Attention Distillation as a Refinement Prior for Compact Object Detectors

YoungJae Cheong, Jhonghyun An

arXiv 2608.18590首次发表:更新:

发表机构

Gachon University(嘉泉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出FD-CanKD框架,通过预测、关系、频域三级知识迁移实现紧凑检测器的高效蒸馏,在COCO数据集上达到48.87 mAP50:95,移除蒸馏模块后学生模型参数为19.7M且性能优于基线。

AI 中文摘要

紧凑目标检测器适用于资源受限的视觉感知任务,但其有限的表征能力导致与大型模型存在精度差距。传统检测器蒸馏通常依赖预测级监督或单一特征对齐目标,如响应、分布、相关性或频域匹配。本文提出面向检测器的框架频率解耦交叉注意力知识蒸馏(FD-CanKD),该框架在三个互补层面迁移教师知识:头部级预测监督、关系级非局部上下文迁移以及频域分量选择性对齐。学生特征首先通过基于交叉注意力的关系迁移聚合教师侧空间上下文,随后频域感知对齐保留互补的结构与细节敏感线索。在受控的微软通用对象上下文(COCO)实验中,固定50个epoch的从零开始训练对比显示,FD-CanKD与代表性检测器知识蒸馏基线相比仍具竞争力;蒸馏后继续微调可生成比仅检测器微调更强的可精细化学生模型,经20个额外epoch后,在交并比阈值0.50至0.95(mAP50:95)下达到48.87平均精度均值(mAP)、65.84 mAP50及53.40 mAP75。训练后移除所有蒸馏模块,部署的学生模型参数保持19.7M不变。该框架在受控的YOLOv12师生设置中作为代表性紧凑检测器案例研究进行实例化与评估。

英文摘要

Compact object detectors are suitable for resource-constrained visual perception, but their limited representation capacity creates an accuracy gap relative to large models. Conventional detector distillation often relies on prediction-level supervision or a single feature-alignment target, such as response, distribution, correlation, or frequency-domain matching. Frequency-Decoupled Cross-Attention Knowledge Distillation (FD-CanKD) is presented as a detector-oriented framework that transfers teacher knowledge at three complementary levels: head-level prediction supervision, relation-level non-local context transfer, and frequency-level component-selective alignment. Student features first aggregate teacher-side spatial context through cross-attention-based relation transfer, after which frequency-aware alignment preserves complementary structural and detail-sensitive cues. Under controlled Microsoft Common Objects in Context (COCO) experiments, fixed 50-epoch from-scratch comparisons show that FD-CanKD remains competitive with representative detector knowledge distillation baselines. Post-distillation continued fine-tuning further produces a stronger refinement-ready student than detector-only fine-tuning, reaching 48.87 mean average precision (mAP) at intersection-over-union thresholds from 0.50 to 0.95 (mAP50:95), 65.84 mAP50, and 53.40 mAP75 after 20 additional epochs. All distillation modules are removed after training, leaving the deployed student unchanged at 19.7M parameters. The framework is instantiated and evaluated in a controlled YOLOv12 teacher-student setting as a representative compact-detector case study.

Comments16 pages, 5 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑