arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SPARK-SAM:面向红外小目标分割SAM的、基于响应知识的自提示适配

SPARK-SAM: Learning How to Prompt and Respond for Infrared Small Target Segmentation

Aji Mao, Zhenming Peng, Bailin Mu, Tian Pu

arXiv 2608.20754首次发表:更新:

AI 中文总结

针对SAM直接迁移至红外小目标分割存在的提示-响应不匹配问题,提出SPARK-SAM方法,仅增0.726M参数即实现高IoU,在两个基准上排名第一。

AI 中文摘要

可提示分割模型提供了可复用的接口,但直接迁移至自动红外小目标分割(IRSTD)时,会出现空间提示与目标域掩码响应不匹配的问题。在使用从测试参考掩码确定性导出的、覆盖目标的宽松框提示进行诊断时,最优官方SAM2.1在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上的IoU仅为4.69%、1.64%和2.28%。本文提出SPARK-SAM(Self-Prompt Adaptation with Response Knowledge for SAM),该模型学习目标域响应知识,并通过图像条件联合自提示状态对解码器进行条件设置。训练过程结合了基准掩码监督与感知可靠性的响应引导。SPARK-SAM仅增加0.726M额外参数,即可在NUAA-SIRST、NUDT-SIRST和IRSTD-1K上达到75.78%、86.49%和68.34%的IoU,在14种重新训练的SAM变体及适配模型中,作为自动图像到掩码方法评估时,在两个基准上排名第一。分阶段IRSTD-1K诊断显示,在预测点获得可靠目标定位前,响应适配已达到最终IoU的大部分。提示监督使预测的提示候选与目标位置对齐,冻结权重干预可测量输出对联合自提示状态的敏感性。匹配的 ablation 实验表明,响应引导和高分辨率提示细化在所有三个数据集上均带来一致的精度提升。代码可在此https URL获取。

英文摘要

Promptable segmentation models provide a reusable interface, but direct transfer to automatic infrared small-target segmentation (IRSTD) exposes a mismatch between spatial prompts and target-domain mask responses. In a diagnostic using target-covering loose-box prompts deterministically derived from test reference masks, the best official SAM2.1 results are only 4.69%, 1.64%, and 2.28% IoU on NUAA-SIRST, NUDT-SIRST, and IRSTD-1K. We introduce SPARK-SAM (Self-Prompt Adaptation with Response Knowledge for SAM), which learns target-domain response knowledge and conditions the decoder through an image-conditioned joint self-prompt state. Training combines benchmark-mask supervision with reliability-aware response guidance. SPARK-SAM achieves 75.78%, 86.49%, and 68.34% IoU with 0.726M additional parameters, ranking first on two benchmarks among 14 retrained SAM variants and adaptations evaluated as automatic image-to-mask methods. The staged IRSTD-1K diagnostic shows that response adaptation reaches most of the final IoU before the predicted points acquire reliable target grounding. Prompt supervision aligns the predicted prompt candidates with target locations, and frozen-weight interventions measure output sensitivity to the joint self-prompt state. Matched ablations show consistent accuracy gains from response guidance and high-resolution prompt refinement across all three datasets. Code is available at https://github.com/Sakauma/SPARK-SAM.

Comments9 pages, 5 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑