arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00666cs.CV

DGNet:用于红外小目标检测的双知识引导网络

DGNet: Dual-knowledge Guided Network for Infrared Small Target Detection

Chenglong Yu, Mingzhu Xu, Jing Wang, Tongtong Wang, Pingping Miao, Liqiang Nie

首次发表
浏览论文内容

中文总结 AI 辅助

针对红外小目标检测中文本引导方法的语义纠缠与部署限制问题,提出双知识引导网络DGNet,通过PWM模块与CDA损失实现性能提升,在三个公开数据集上验证了其有效性。

中文摘要 AI 辅助

红外小目标检测(Infrared Small Target Detection, IRSTD)是计算机视觉领域一项重要且具有挑战性的任务。近年来,文本引导方法显著提升了检测性能,但仍存在两个关键局限:其一,单一文本描述同时建模背景与目标会导致语义纠缠,与背景抑制、目标增强的目标相悖;其二,依赖特定图像的文本提示(推理时需额外使用CLIP等外部模型)造成部署限制。为解决这些问题,本文提出一种基于多种可泛化文本的新型双知识引导网络(Dual-knowledge Guided Network, DGNet)。具体而言,设计了先验知识小波调制(Prior-knowledge Wavelet Modulation, PWM)模块,利用分别表征大规模背景与稀疏目标的双文本先验,在频域有效解缠并调制纠缠语义;还引入了共识知识方向对齐(Consensus-knowledge Directional Alignment, CDA)损失,将样本的初始状态与理想目标分别建模为“复杂背景”和“明亮目标”,从而为模型构建清晰统一的方向优化轨迹。在三个公开数据集上开展的大量实验,证明了DGNet的优越性能及各组件的有效性,源代码可访问此https URL。

英文摘要

InfRared Small Target Detection (IRSTD) is a prominent and challenging task in computer vision. In recent years, text-guided methods have significantly improved detection performance. However, they still suffer from two key limitations. First, a single text description simultaneously modeling both background and target leads to semantic entanglement, which contradicts the objective of background suppression and target enhancement. Second, reliance on image-specific textual prompts (requiring additional external models such as CLIP during inference) results in deployment constraints. To address these issues, we propose a novel Dual-knowledge Guided Network (DGNet) based on multiple generalizable texts. Specifically, we design a Prior-knowledge Wavelet Modulation (PWM) module, which leverages dual textual priors that separately characterize large-scale backgrounds and sparse targets to effectively disentangle and modulate entangled semantics in the frequency domain. Furthermore, we introduce a Consensus-knowledge Directional Alignment (CDA) loss, which models the initial state and the ideal target across samples as `complex background' and `bright target', respectively, thereby constructing a clear and unified directional optimization trajectory for the model. Extensive experiments on three public datasets demonstrate the superior performance of DGNet and the effectiveness of each component. The source code is available at https://github.com/iLearn-Lab/MM26-DGNet.

发表机构

  • Shandong University(山东大学)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑