利用MLLMs的目标知识实现鲁棒的少样本分割
Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation
- UCAS(中国科学院大学)
- ISCAS(中国科学院软件研究所)
- UNT(北德克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出MK-FSS框架,利用多模态大语言模型挖掘查询图像的空间与语义目标知识,通过双记忆融合和跨模态提示增强SAM 2的少样本分割鲁棒性,显著优于现有方法。
AI中文摘要:
少样本分割(FSS)旨在利用少量(例如一个或五个)标注示例来分割未见过的对象类别,从而实现对新颖类别的有效适应。传统模型通常依赖支持图像与查询图像之间基于外观的视觉匹配来进行分割。尽管这些方法直接明了,但由于缺乏足够的目标知识,它们常常难以处理查询图像中显著的外观差异和遮挡问题。为缓解这一问题,我们引入了一种新颖框架,利用多模态大语言模型(MLLMs)强大的推理能力来挖掘目标知识,并将其用于增强FSS。具体而言,基于SAM 2,我们提出的方法MK-FSS利用MLLM从查询图像中提取的两种互补知识进行FSS,包括提供指示潜在目标位置的空间先验的空间知识,以及用文本描述目标的语义知识。空间知识首先被编码为记忆表示,然后通过精心设计的双记忆辩论融合(DMDF)模块,将所得记忆与来自查询图像的支持引导记忆特征相集成,从而产生更鲁棒的目标记忆特征。与此同时,语义知识被编码为文本特征,并通过渐进式跨模态提示生成器(PCPG)与多尺度查询特征融合,生成用于分割的目标感知多模态提示。通过协同工作,双记忆特征和多模态提示提供了对目标的全面表示,从而实现更鲁棒的分割。在我们广泛的实验中,MK-FSS展示了有前景的结果,并大幅超越了现有方法。代码将发布。
英文摘要:
Few-shot segmentation (FSS) aims to segment unseen object categories with a few (e.g., one or five) labeled examples, enabling efficient adaptation to novel classes. Conventional models typically rely on appearance-based visual matching between support and query images for segmentation. While straightforward, these methods often struggle to handle significant appearance discrepancies and occlusions in the query image due to insufficient target knowledge. To mitigate this, we introduce a novel framework that mines target knowledge using the strong reasoning capacity of Multimodal Large Language Models (MLLMs) and employs it to enhance FSS. Specifically, building on SAM 2, our method, named MK-FSS, exploits two forms of complementary knowledge derived from a query image by an MLLM for FSS, including spatial knowledge, which provides a spatial prior indicating the potential target location, and semantic knowledge, which describes the target using text. The spatial knowledge is first encoded into a memory representation, and then resulting memory is integrated with the support-guided memory feature from query image through a carefully designed dual-memory debate-fusion (DMDF) module, yielding a more robust target memory feature. In parallel, the semantic knowledge is encoded into the textual feature, which is fused with multi-scale query features via a progressive cross-modal prompt generator (PCPG), producing a target-aware multimodal prompt for segmentation. Working together, the dual-memory feature and the multimodal prompt provide a comprehensive representation of the target, enabling more robust segmentation. In our extensive experiments, MK-FSS shows promising results and largely surpasses existing methods. Code will be released.