对齐共识教学:弱对齐可见光-红外图像中标签高效的有向目标检测
Aligned Consensus Teaching for Label-Efficient Oriented Object Detection in Weakly-Aligned Visible-Infrared Imagery
浏览论文内容
中文总结 AI 辅助
针对弱对齐可见光-红外图像中标签高效的有向目标检测,提出对齐共识教师(ACT)框架,通过循环一致区域对齐、跨模态共识均值教师和文本引导增强,在10%标注下达到全监督mAP的94.3%。
中文摘要 AI 辅助
可见光-红外目标检测(VIOD)从配对的可见光和红外图像中检测有向边界框目标。现有方法依赖成本高昂的双模态标注。半监督学习可以减轻这一负担,但将其从单模态检测扩展到VIOD具有挑战性。在本文考虑的实际图像对级别设置中,仅少数图像对在两种模态下都有标注,而其余图像对完全未标注。这种有限的监督带来了三个挑战:(i)标注框太少,无法实现稳健的跨模态对齐;(ii)自训练过程中,由分支遗漏导致的伪标签错误会不断累积;(iii)随着标注预算减少,尾部类别标注变得极其稀缺。我们提出了对齐共识教师(ACT)框架,用于在此设置下进行标签高效的VIOD。其循环一致区域对齐(CRA)结合了循环一致性和稀疏锚点与可靠性加权的区域匹配。跨模态共识均值教师(CMC-MT)在保持图像对一致的视图下形成共识伪标签,以恢复分支遗漏并监督未标注的图像对。文本引导的跨模态实例增强(TG-CMIA)利用视觉-语言场景先验来合成尾部类别实例对,同时保留RGB-IR偏移。据我们所知,ACT是首个在此图像对级别设置下研究半监督VIOD的框架。在DroneVehicle和VEDAI上的实验表明,在不同标注比例下均有一致的性能提升。在DroneVehicle上使用10%的标注图像对时,ACT达到了相同检测器在全监督下获得的mAP的94.3%。代码和模型将在GitHub上发布,以促进未来的研究。
英文摘要
Visible-infrared object detection (VIOD) detects objects with oriented bounding boxes from paired visible and infrared images. Existing methods depend on costly dual-modality annotations. Semi-supervised learning can reduce this burden, but extending it from single-modal detection to VIOD is challenging. In the practical image-pair-level setting considered here, only a few pairs are labeled in both modalities, while the rest are completely unlabeled. This limited supervision creates three challenges: (i) too few labeled boxes for robust cross-modal alignment; (ii) pseudo-label errors caused by branch-wise misses accumulate during self-training; and (iii) tail-class annotations become critically scarce as the labeling budget decreases. We propose Aligned Consensus Teacher (ACT) for label-efficient VIOD in this setting. Its Cycle-Consistent Region Alignment (CRA) combines cycle consistency and sparse anchors with reliability-weighted regional matching. Cross-Modal Consensus Mean-Teacher (CMC-MT) forms consensus pseudo labels under pair-preserving views to recover branch-wise misses and supervise unlabeled pairs. Text-Guided Cross-Modal Instance Augmentation (TG-CMIA) uses a vision-language scene prior to compose tail-class instance pairs while preserving RGB--IR offsets. To the best of our knowledge, ACT is the first framework to study semi-supervised VIOD under this image-pair-level setting. Experiments on DroneVehicle and VEDAI show consistent gains across annotation ratios. With 10\% labeled pairs on DroneVehicle, ACT reaches 94.3\% of the mAP obtained by the same detector under full supervision. Code and models will be available on GitHub to facilitate future work.
发表机构
- Beijing University of Technology(北京工业大学)
- University of Science and Technology of China(中国科学技术大学)
- Peking University(北京大学)
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- Beijing Institute of Technology(北京理工大学)
- Fuzhou University(福州大学)
- China University of Geosciences Wuhan(中国地质大学(武汉))
- Ghent University(根特大学)
机构由 AI 辅助整理,请以论文原文为准。