CMAMBADEPTH:基于通道Mamba与混合注意力的自监督单目深度估计
CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention
另 1 家 · 查看机构详情
- Harbin Engineering University(哈尔滨工程大学)
- Key Laboratory of Advanced Marine Communication and Information Technology(先进海洋通信与信息技术重点实验室)
- University of Toronto(多伦多大学)
- Kanagawa University(神奈川大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出CMambaDepth,一种基于通道Mamba与混合注意力的自监督单目深度估计框架,通过双向与单向通道Mamba实现高效多尺度特征融合,并引入混合注意力模块平衡局部与全局建模,在KITTI、DDAD及NYUv2上取得优异性能。
中文摘要 AI 辅助
精确的单目深度估计是单相机场景理解的核心使能技术。然而,现有的自监督单目深度估计方法普遍面临跨尺度信息交互效率低以及难以平衡局部与全局空间建模的瓶颈。本文提出CMambaDepth,一种自监督框架,通过通道选择性状态传播实现高效的多尺度特征融合和细粒度上下文建模。具体而言,双向通道Mamba(Bi-CMamba)跨尺度对齐编码器特征,并在有序尺度组之间实现双向信息交换。单向通道Mamba(Uni-CMamba)逐步聚合解码器特征,并通过组选择机制保留细粒度尺度组以供后续融合。此外,引入混合注意力模块(HAM),将大核局部上下文与曼哈顿自注意力相结合,实现互补的空间建模。实验结果表明,我们的方法取得了极具竞争力的性能。具体而言,我们的模型在KITTI上达到AbsRel为0.094、RMSE为4.156,在DDAD上AbsRel为0.140。在NYUv2上的零样本跨数据集泛化测试中,其AbsRel达到0.232,优于基线RA-Depth达7.2%。
英文摘要
Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and global spatial modeling. In this paper, we propose CMambaDepth, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation. Specifically, Bidirectional Channel Mamba (Bi-CMamba) aligns encoder features across scales and enables bidirectional information exchange among ordered scale groups. Unidirectional Channel Mamba (Uni-CMamba) progressively aggregates decoder features and retains fine-grained scale groups through a group selection mechanism for subsequent fusion. Furthermore, a Hybrid Attention Module (HAM) is introduced to combine large-kernel local context and Manhattan self-attention for complementary spatial modeling. Experimental results demonstrate that our method achieves highly competitive performance. Specifically, our model achieves an AbsRel of 0.094 and an RMSE of 4.156 on KITTI, and an AbsRel of 0.140 on DDAD. In the zero-shot cross-dataset generalization test on NYUv2, it attains an AbsRel of 0.232, outperforming the baseline RA-Depth by 7.2%.