提示与精炼:用于带噪声标签的红外小目标检测的非对称互学习
Prompt and Refinement: Asymmetric Mutual Learning for Infrared Small Target Detection with Noisy Labels
浏览论文内容
中文总结 AI 辅助
针对红外小目标检测中噪声标签问题,提出非对称互学习范式PAR,结合SAM与专用检测器协同精炼标签,实现最先进性能。
中文摘要 AI 辅助
现有的基于数据驱动的红外小目标检测(ISTD)方法通常需要大规模数据集以及精确的像素级标注来训练模型。然而,在现实应用中,由于对专家知识的重度依赖以及红外小目标本身固有的弱显著性,这种劳动密集型的需求难以满足。因此,在模型训练过程中不可避免地会出现噪声标签,这会严重误导目标感知的学习,使其趋向于虚假的模式。为了解决这一挑战,我们提出了提示与精炼(PAR),一种针对ISTD的标签噪声鲁棒的非对称互学习范式。具体来说,PAR包含一个预训练的Segment Anything Model(SAM)和一个从头训练的ISTD专用检测器,它们通过一种同伴教学方案协同学习。结合局部对比度规律,两个非对称同伴模型的预测被相互利用作为对方监督掩码的修正线索。互补的归纳偏置之间的交互有效地防止了标签修正过程退化为单一模型的自我确认循环,从而使得标注能够朝着内在目标特征逐步精炼。此外,检测器的预测被用作修正掩码提示,以促进视觉基础模型的任务特定适应。同时,在优化过程中引入了一种证据不确定性估计策略,以进一步减轻噪声标签的不利影响。在三个ISTD数据集上进行的多种噪声标签场景下的广泛实验表明,PAR始终达到了最先进的性能。
英文摘要
Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications due to the heavy reliance on expert knowledge and the inherently weak distinctiveness of infrared small targets. Consequently, the presence of noisy labels during model training is inevitable, which can severely mislead the learning of target perception toward spurious patterns. To address this challenge, we propose Prompt and Refinement (PAR), a label-noise-robust asymmetric mutual learning paradigm for ISTD. Specifically, PAR comprises a pretrained Segment Anything Model (SAM) and an ISTD-specific detector trained from scratch, which learn collaboratively through a peer-teaching scheme. Coupled with local contrast regularity, the predictions of the two asymmetric peer models are mutually exploited as rectification cues for the supervisory masks of their counterparts. The interaction between complementary inductive biases effectively prevents the label correction process from degenerating into the self-confirmation loop of a single model, enabling progressive refinement of the annotations toward intrinsic target characteristics. In addition, the detector predictions are utilized as corrective mask prompts to facilitate task-specific adaptation of the vision foundation model. Moreover, an evidential uncertainty estimation strategy is introduced into the optimization process to further alleviate the adverse effects of noisy labels. Extensive experiments under diverse noisy label scenarios on three ISTD datasets demonstrate that PAR consistently achieves state-of-the-art performance.
发表机构
- Hong Kong Baptist University(香港浸会大学)
- Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。