发表机构
National University of Defense Technology; College of Computer Science, Nankai University; Peking University(国防科技大学; 南开大学计算机学院; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全域红外小目标检测的域偏移问题,提出“先理解再检测”范式及JinSight方法,构建视觉-语言数据集OmniIRST-VL,实现跨异构红外域的泛化检测。
AI 中文摘要
全域红外小目标(IRST)检测对红外监视至关重要,但由于异构成像域和不一致的目标特性,该任务仍具挑战性。此前基于深度学习的方法针对仅视觉范式开发,在特定域任务上取得了良好性能,但现有方法遵循特定任务的监督学习范式,该范式将全场景红外观测简化为稀疏目标监督,丢弃了异构域间保持不变的语义,导致域偏移下检测性能大幅下降。为解决该问题,本文提出“先理解再检测”范式,将全域IRST检测构建为理解驱动的过程,即先进行整体红外目标理解,再开展精确检测。基于该范式,本文提出JinSight方法,该方法首先通过语言监督开发整体IRST理解,再将学习到的跨域表征迁移至精确小目标检测;通过将红外表征与语言语义建立关联,JinSight使单个模型能够在异构红外域间泛化。此外,本文提出潜在语义交互(LSI)模块,该模块在紧凑低秩空间中交换与语言对齐的全局语义和细粒度空间特征。为解决多模态全域IRST基准缺失的问题,本文构建了首个大规模、高多样性的全域IRST检测视觉-语言数据集OmniIRST-VL,该数据集包含6项互补指令任务的超3.9万条标注,涵盖场景级理解和以目标为中心的推理。
英文摘要
Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target characteristics. Previous deep learning-based methods have been developed for visual-only paradigms and achieved promising performance on domain-specific tasks. However, existing methods follow the task-specific supervised learning paradigm. This paradigm simplifies the full-scene infrared observations to sparse target supervision, discarding the semantics that remain invariant across heterogeneous domains. Consequently, detection performance suffers substantially under domain shifts. To handle this issue, we introduce \textbf{``understand before detect''}, a paradigm that formulates omni-domain IRST detection as an understanding-driven process, where holistic infrared target understanding precedes precise detection. Building on this paradigm, we propose \textbf{JinSight}, which first develops holistic IRST understanding through language supervision and then transfers the learned cross-domain representations to precise small-target detection. By grounding infrared representations in language semantics, JinSight enables a single model to generalize across heterogeneous infrared domains. We then introduce Latent Semantic Interaction (LSI), which exchanges language-aligned global semantics with fine-grained spatial features in a compact low-rank space. To address the lack of multimodal omni-domain IRST benchmarks, we build \textbf{OmniIRST-VL}, the first large-scale, highly diverse vision--language dataset for omni-domain IRST detection. It comprises over 39k annotations across six complementary instruction tasks covering both scene-level understanding and target-centric reasoning.