发表机构
Korea Institute of Science and Technology; Center for Humanoid Research Korea Institute of Science and Technology(韩国科学技术院; 韩国科学技术院类人机器人研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究视觉曼巴和MambaOut在表示视觉信息上的差异,通过交叉模型中心核对齐分析等方法,发现二者编码策略不同,视觉曼巴在密集预测中有优势,其优势源于语义证据在令牌幅度和方向上的组织方式,揭示令牌幅度和方向结构对改进视觉主干的关键作用。
AI 中文摘要
视觉曼巴模型用线性复杂度的选择性状态空间模型(SSM)取代了二次自注意力,成为高效的视觉主干。然而,MambaOut表明门控卷积神经网络块在图像分类上能与视觉曼巴匹配或超越,质疑了SSM在视觉中的必要性。本文通过交叉模型中心核对齐分析发现视觉曼巴最后阶段块形成的表示与MambaOut及其自身前序块明显不同。将每个空间令牌分解为幅度和方向后发现,MambaOut将类判别信息集中在与Grad-CAM归因对齐的高范数前景令牌中,而视觉曼巴主要在背景区域产生高范数令牌,与Grad-CAM不对齐,但在令牌方向上保留判别信号。研究还将这种差异与高分辨率分类和语义分割联系起来,在全量微调分割时,视觉曼巴始终优于MambaOut。结果表明视觉曼巴在密集预测中的优势不仅源于SSM机制或序列长度,还源于语义证据在令牌幅度和方向上的组织方式。最终得出令牌幅度和方向结构是改进视觉主干的关键轴,特别是在密集监督下。
英文摘要
Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones. However, MambaOut demonstrates that a Gated CNN block can match or exceed VMamba on image classification, questioning the necessity of SSMs for vision. This raises a fundamental question: do VMamba and MambaOut encode visual information differently at the representation level? To investigate, we apply cross model centered kernel alignment (CKA) analysis and find that VMamba's final stage blocks form representations distinctly different from both MambaOut and its own preceding blocks. We therefore focus on the final block features, decomposing each spatial token into magnitude and direction. MambaOut concentrates class-discriminative information in high-norm foreground tokens that align with Grad-CAM attribution. VMamba, by contrast, produces high-norm tokens predominantly in background regions, misaligned with Grad-CAM, yet preserves discriminative signals primarily in token directions. These observations reveal that the two models rely on different encoding strategies. We connect this difference to high-resolution classification and semantic segmentation. VMamba distributes logit support broadly across object regions, whereas MambaOut relies on sparse dominant tokens, a strategy that becomes less stable as token counts grow. Under full fine-tuning for segmentation, VMamba consistently outperforms MambaOut. These results suggest that VMamba's advantage in dense prediction stems not merely from the SSM mechanism or sequence length, but from how semantic evidence is organized across token magnitude, direction. Ultimately, we conclude that token magnitude and directional structure serve as critical axes for improving visual backbones, particularly under dense supervision.