arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IRSTD-Agent:基于缩放引导交互学习的智能体红外小目标检测

IRSTD-Agent: Agentic Infrared Small Target Detection via Zoom-Guided Interaction Learning

Jiawen Xi, Yu Zhang, Tianyi Zhao, Zhu Liu, Maoxun Yuan, Xingxing Wei

arXiv 2610.05342首次发表:更新:

发表机构

Beihang University; Dalian University of Technology(北京航空航天大学; 大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出IRSTD-Agent,一种基于缩放引导交互学习的智能体框架,通过动态视觉搜索和五种互补工具实现红外小目标精确检测,在WideIRSTD-Full和IRSTD-1k数据集上优于现有视觉语言模型。

AI 中文摘要

红外小目标检测在海上监视和空中侦察中扮演着重要角色。尽管多模态大语言模型(MLLMs)为视觉理解提供了有前景的能力,但现有的基于MLLM的方法难以精确地定位红外小目标。在本文中,我们提出了IRSTD-Agent,一种通过动态视觉搜索进行红外小目标检测的智能体框架。该框架使MLLM能够自适应地决定在何处以及以何种尺度检查图像,并逐步收集细粒度的视觉证据以实现精确的目标定位。五种互补的视觉工具(PROPOSAL、ZOOM、DETECT、DROP和REFINE)支持目标候选发现、自适应观察、目标定位、假设拒绝和目标范围细化,共同在原始分辨率图像上实现协调的搜索过程。为了教会MLLMs执行这种搜索,我们引入了缩放引导交互学习,该方法利用从标注中导出的交互轨迹来监督工具选择及其相应参数。通过在WideIRSTD-Full和IRSTD-1k数据集上的广泛实验,我们证明了IRSTD-Agent优于所评估的视觉语言模型,并增强了MLLMs在IRSTD任务中的精确目标定位能力。

英文摘要

Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose IRSTD-Agent, an agentic framework for infrared small target detection through dynamic visual search. The framework enables an MLLM to adaptively determine where and at what scale to inspect an image and progressively gather fine-grained visual evidence for precise target localization. Five complementary visual tools (PROPOSAL, ZOOM, DETECT, DROP and REFINE) support object candidate discovery, adaptive observation, target localization, hypothesis rejection, and target extent refinement, together enabling a coordinated search process over original-resolution images. To teach the MLLMs to conduct this search, we introduce Zoom-guided Interaction Learning, which uses annotation-derived interaction trajectories to supervise tool selection and the corresponding arguments. Through extensive experiments on WideIRSTD-Full and IRSTD-1k datasets, we demonstrate that IRSTD-Agent outperforms the evaluated vision-language models and enhances the precise localization capabilities of MLLMs in IRSTD tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑