LAD-COD:面向伪装目标检测的语言对齐密集感知
LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection
浏览论文内容
中文总结 AI 辅助
本研究针对伪装目标检测中密集视觉特征缺乏语言指导的问题,提出LAD-COD框架,通过语言对齐双视觉融合实现语义与分层视觉特征的对齐,在三个数据集的12组对比中取得最优结果。
中文摘要 AI 辅助
伪装目标检测(COD)旨在分割与周围环境视觉相似度极高的物体,这会降低前景与背景的可区分性,并削弱外观、纹理和结构层面的边界证据。这类局限性促使人们采用指令条件语义作为自上而下的指导,以识别哪些微弱视觉线索与目标相关。近期基于大型多模态模型(LMM)构建的分割系统,通过指令条件目标嵌入来指导掩码解码,证明了这种可能性。然而,在这种从语言到掩码的范式中,生成的目标嵌入主要作用于掩码解码器,而必须保留低对比度边界和精细局部结构的密集视觉特征却缺乏显式指导。我们提出了面向COD的语言对齐密集感知框架LAD-COD,该框架将自上而下的语义目标指导与自下而上的分层视觉特征进行对齐。LAD-COD未完全适配大型通用图像编码器,而是学习了一个可训练的分层视觉分支,用于捕捉对伪装敏感的纹理、边界和上下文信息。为将这些特征与目标嵌入对齐,LAD-COD采用了语言对齐双视觉融合(LADVF),该方法将嵌入从稀疏提示扩展为查询补丁级语言对齐特征,并对其与分层特征的残差集成进行门控。这种设计使语义信息能够指导定位,同时保留伪装分割所需的精细结构细节。在CAMO、COD10K和NC4K数据集上开展的实验表明,LAD-COD在全部12组数据集-指标对比中均取得了报告的最优值。
英文摘要
Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as top-down guidance for identifying which weak visual cues are relevant to the target. Recent segmentation systems built on large multimodal models (LMMs) demonstrate this possibility through instruction-conditioned target embeddings that guide mask decoding. However, in this language-to-mask paradigm, the generated target embedding conditions mainly the mask decoder, leaving the dense visual features that must preserve low-contrast boundaries and fine local structure without explicit guidance. We propose Language-Aligned Dense perception for COD (LAD-COD), a framework that aligns top-down semantic target guidance with bottom-up hierarchical visual features. Instead of fully adapting a large generic image encoder, LAD-COD learns a trainable hierarchical visual branch that captures camouflage-sensitive texture, boundary, and contextual information. To align these features with the target embedding, LAD-COD applies Language-Aligned Dual Visual Fusion (LADVF), which extends the embedding beyond sparse prompting to query patch-level language-aligned features and to gate their residual integration with the hierarchical features. This design allows semantic information to guide localization while preserving the fine structural details needed for camouflage segmentation. Experiments on CAMO, COD10K, and NC4K show that LAD-COD obtains the best reported value in all 12 dataset-metric comparisons.