发表机构
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,教育部,厦门大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究指认伪装物体分割问题,提出GFR-SAM三阶段免训练框架,通过生成候选掩码、过滤背景干扰、细化边界等步骤,在R2C7K基准测试中表现出色,弥合通用模型与特定感知任务差距,凸显SAM3跨图像提示潜力。
AI 中文摘要
指认伪装物体检测(Ref-COD)需要在参考线索引导下分割隐藏目标。有监督方法标注量大,基于稀疏点提示的免训练方法对定位误差敏感。我们提出了GFR-SAM,一个强大的三阶段免训练框架。它将范式从脆弱的点匹配转变为“生成-过滤-细化”流程。首先通过上下文示例引导分割,利用跨图像推理生成候选掩码;接着通过基于DINOv3的原型对齐对候选者进行排序以抑制背景干扰;最后通过几何语义细化模块协同边界框和文本提示恢复细粒度边界并提高实例召回率。在R2C7K基准测试中,GFR-SAM在加权F-measure上比现有免训练方法高出8.7%,与有监督的最先进方法竞争。这项工作强调了释放SAM3跨图像上下文提示潜在能力的潜力,建立了一个强大的免训练范式,无需特定任务微调就能有效弥合通用基础模型与专门的、标签密集型感知任务之间的差距。
英文摘要
Referring Camouflaged Object Detection (Ref-COD) requires segmenting hidden targets guided by reference cues. While supervised methods are annotation-heavy and training-free approaches via sparse point-prompting are sensitive to localization errors, we propose GFR-SAM, a robust three-stage training-free framework. GFR-SAM shifts the paradigm from fragile point-matching to a "Generate-Filter-Refine" pipeline. First, we introduce In-Context Exemplar-guided Segmentation, empowering SAM3 with cross-image inference to generate candidate masks via holistic visual exemplars, bypassing its native intra-image constraints. Second, a Region-Global Contrastive Filtering module ranks candidates through DINOv3-based prototypical alignment, effectively suppressing background distractors. Finally, a Geometric-Semantic Refinement module synergizes bounding box and text prompts to recover fine-grained boundaries and enhance instance recall. Evaluated on the R2C7K benchmark, GFR-SAM outperforms existing training-free methods by 8.7\% in weighted F-measure ($F_β^w$) and competes with supervised state-of-the-art counterparts. Ultimately, this work underscores the potential of unlocking SAM3's latent capability for cross-image In-Context prompting, establishing a robust, training-free paradigm that effectively bridges the gap between general-purpose foundation models and specialized, label-intensive perception tasks without the need for task-specific fine-tuning.