发表机构
Taizhou Institute of Science and Technology, Nanjing University of Science and Technology; University of Arizona(南京理工大学泰州科技学院; 亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出PCLM方法,利用冻结CLIP编码器,通过原型对比和局部放大实现小目标定位,在多个数据集上显著优于现有基线。
AI 中文摘要
小目标在视觉-语言编码器中仅占据少量图像块,因此其空间特征常常将物体外观与周围内容混合在一起。我们提出了原型对比与局部放大(PCLM),一种利用冻结的CLIP编码器的支持条件化定位方法。每个类别使用五张掩蔽支持图像,通过等权重的区域特征定义前景和背景原型。两者的差异为查询图像块提供了共享的评分方向,明确地将目标证据与所展示的背景进行比较。九个重叠的查询窗口被放大并独立编码,以更密集地采样小目标。重投影和覆盖率平均将它们的得分合并为连续的定位图。类别方向仅占用2 KiB,与支持样本数量无关,并且可以在具有映射类别的数据集之间不变地迁移。在来自VOC、COCO、ADE20K和Oxford-IIIT Pets的5,047个小目标查询上,在我们的评估协议下,PCLM在每个数据集上的平均像素AP均高于所有评估的文本条件定位基线。相对于最强的场景数据集基线,增益范围为5.63至13.23个百分点。在可比较的测量延迟下,局部放大相对于整幅画布放大,将场景小目标AP提高了4.51至5.68个百分点。因子实验表明,原型对比增强了局部观察的益处,包括在匹配的图像坐标过滤条件下。支持预算实验表明,额外的示例可以细化类别估计,而不会增加表示大小或查询时评分成本。
英文摘要
Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.