arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27843cs.CV

VCP-DCN:超越视觉隐蔽属性的深度协作网络用于伪装目标检测

VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

Songsong Duan, Xi Yang, Nannan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有伪装目标检测方法忽略深度域隐蔽目标模态特性的问题,提出VCP-DCN网络,通过SPE、MDA、DAI模块逐步实现多模态对齐、交互与融合,在三个权威数据集上验证了有效性。

中文摘要 AI 辅助

伪装目标检测(COD)旨在识别并分割复杂环境中的伪装目标,这类目标因颜色和纹理与背景相似而常被隐蔽。现有若干COD方法引入深度图,通过学习互补的RGB-D特征提升检测性能,但忽略了深度域中隐蔽目标的模态特异性特征。为解决该问题,我们提出名为VCP-DCN的深度协作网络,用于在深度域中挖掘超越视觉隐蔽原型的可区分多模态特征。具体而言,VCP-DCN针对COD任务逐步执行多模态对齐、交互与融合:在对齐阶段,我们提出可分离原型嵌入(SPE)模块,通过原型对比学习学习模态一致性与模态特异性的RGB/深度原型令牌;在交互阶段,我们开发多模态双注意力(MDA)模块,通过模态一致性RGB/深度原型令牌与视觉令牌间的局部响应图增强跨模态特征表示;在融合阶段,我们设计深度自适应注入(DAI)模块,采用决策机制自适应衡量RGB/深度特征的贡献,该机制计算RGB/深度模态特异性原型令牌与模态一致性原型令牌间的相似距离。大量实验表明,我们的VCP-DCN在三个权威数据集上具有有效性。

英文摘要

Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth domain. To address this issue, we propose a depth collaborative network, called VCP-DCN, to mine distinguishable multi-modality features beyond visual concealed prototype in depth domain. Specifically, VCP-DCN progressively performs multi-modality alignment, interaction, and fusion for the COD task. In the \textbf{alignment} stage, we propose a Separable Prototype Embedding (SPE) module to learn modality-consistency and modality-specific RGB/depth prototype tokens through prototype contrastive learning. Furthermore, we develop a Multi-modality Dual Attention (MDA) module to enhance the cross-modal feature representation through local response maps between modality-consistency RGB/depth prototype tokens and visual tokens on the \textbf{interaction} stage. Finally, we design a Depth Adaptive Injection (DAI) module to adaptively measure contribution of RGB/depth features with a decision-making mechanism, which calculates similarity distance between RGB/depth modality-specific prototype tokens and modality-consistency ones on the \textbf{fusion} stage. Extensive experiments demonstrate the effectiveness of our VCP-DCN on three authoritative datasets.

发表机构

  • Xidian University(西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑