发表机构
Shandong University; Harbin Institute of Technology, Shenzhen(山东大学; 哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对红外小目标检测中纯视觉方法难区分目标与杂波、现有多模态方法存在缺陷的问题,提出ADGNet,设计ADP、ADBI、AFA模块,构建AITIR数据集,性能优于21种SOTA方法。
AI 中文摘要
红外小目标检测(IRSTD)是一项具有挑战性的任务。仅依赖像素级信息的纯视觉方法难以区分目标与杂波。当前多模态方法通常用单一文本提示描述目标和背景,该方法缺乏针对性的区域引导,且忽略了红外语义的非对称性,导致背景抑制信息不足,还会引发严重的特征优化冲突,使小目标被噪声掩盖。为解决这些问题,本文提出了一种新型非对称双文本引导网络(ADGNet)。具体而言,考虑到红外语义的非对称性,我们首先设计了非对称双文本提示(ADP),其包含与图像无关的抽象目标提示和与图像相关的详细背景提示。为利用这些提示,我们引入了非对称双分支交互(ADBI)模块,以各自的文本先验分别引导视觉特征,保护目标免受噪声干扰,同时抑制背景杂波。随后,我们引入了自适应特征聚合(AFA)模块,以动态融合两个分支的特征。此外,我们通过为三个公开数据集(IRSTD-1K、NUDT-SIRST和SIRST)提供非对称文本标注,构建了多模态非对称图像-文本红外(AITIR)数据集。大量实验表明,ADGNet的性能优于21种最先进(SOTA)方法。代码可在该https URL获取。
英文摘要
InfRared Small Target Detection (IRSTD) is a challenging task. Relying solely on pixel-level information, vision-only methods struggle to distinguish targets from clutter. Current multimodal methods typically describe both targets and backgrounds with a single textual prompt. Such an approach lacks dedicated regional guidance and ignores infrared semantic asymmetry. Consequently, it provides insufficient background suppression information and introduces severe feature optimization conflicts, overwhelming small targets with noise. To address these issues, we propose a novel Asymmetric Dual-text Guided Network (ADGNet). Specifically, accounting for the infrared semantic asymmetry, we first design the Asymmetric Dual-text Prompt (ADP), comprising an image-agnostic abstract target prompt and an image-specific detailed background prompt. To leverage these prompts, we introduce an Asymmetric Dual-Branch Interaction (ADBI) module to separately guide visual features with their respective text priors, protecting targets from noise while suppressing background clutter. Subsequently, we introduce an Adaptive Feature Aggregation (AFA) module to dynamically fuse features from the two branches. Furthermore, we construct a multimodal Asymmetric Image-Text Infrared (AITIR) dataset by providing asymmetric text annotations for three public datasets (IRSTD-1K, NUDT-SIRST, and SIRST). Extensive experiments demonstrate that ADGNet outperforms 21 state-of-the-art (SOTA) methods. Code is available at https://github.com/iLearn-Lab/MM26-ADGNet.
Comments10 pages, 10 figures. Accepted by the 34th ACM International Conference on Multimedia (ACM Multimedia 2026)