发表机构
Jiangxi Normal University; East China Normal University(江西师范大学; 华东师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ViCo-SAM3通过视觉条件模块动态调制文本嵌入,并设计跨模态绑定模块增强交互,以弥合语义鸿沟,在OVCamo基准上取得最先进性能。
AI 中文摘要
开放词汇伪装目标分割(OVCOS)旨在文本引导下分割未见过的伪装目标。我们观察到,在OVCOS中,SAM3在全局文本语义与细粒度像素级视觉线索之间仍存在显著的语义鸿沟。同时,完全微调文本编码器会引入大量参数开销,并存在对训练类别过拟合的风险,从而损害开放词汇表示的灵活性。为解决这些问题,我们提出了ViCo-SAM3,一种专为OVCOS设计的视觉条件对齐框架。具体而言,我们引入了视觉条件(ViCo)模块,该模块利用全局视觉上下文动态调制文本嵌入,使文本表示能够适应当前图像内容,从而有效弥合视觉与文本之间的语义鸿沟。在此基础上,我们进一步设计了视觉条件跨模态绑定(ViCoBind)模块,以增强视觉与文本表示之间的跨模态交互和语义对齐。无需繁琐的附加组件,ViCo-SAM3在OVCamo基准上取得了最先进的性能,并展现出强大的泛化能力。
英文摘要
Open-vocabulary camouflaged object segmentation (OVCOS) aims to segment unseen camouflaged objects under text guidance. We observe that SAM3 still suffers from a pronounced semantic gap between global textual semantics and fine-grained pixel-level visual cues in OVCOS. Meanwhile, fully fine-tuning the text encoder introduces heavy parameter overhead and risks overfitting to training categories, which compromises open-vocabulary representation flexibility. To address these issues, we propose ViCo-SAM3, a Vision-Conditioned alignment framework designed for OVCOS. Specifically, we introduce vision-conditioned (ViCo) module, which dynamically modulates text embeddings with global visual context, enabling textual representations to adapt to the current image content and thereby effectively bridging the semantic gap between vision and text. Building on this, we further design a vision-conditioned cross-modal binding (ViCoBind) module to enhance cross-modal interaction and semantic alignment between visual and textual representations. Without bells and whistles, ViCo-SAM3 achieves state-of-the-art performance on the OVCamo benchmark and demonstrates strong generalization.