arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24337cs.CV

LiAuto-MindViT:一种具有自适应双向Mamba的混合视觉骨干网络

LiAuto-MindViT: A Hybrid Vision Backbone with Adaptive Bidirectional Mamba

Lifu Mu, Shuai Chen, Wen Zheng, Haoyi Sun, Xueyang Fu, Pengfei Yu, Ning Mao, Tao Wei, Zhou Pan, Kun Zhan

首次发表
浏览论文内容

中文总结 AI 辅助

针对Mamba视觉适配挑战,提出LiAuto-MindViT混合骨干,以自适应双向Mamba消除方向偏差,结合重参数化模块加速推理,在图像分类、检测和分割上达到最优性能。

中文摘要 AI 辅助

虽然基于Mamba的模型在长序列建模方面展现出强大的潜力,但由于视觉理解需要局部邻域相关性和多方向空间上下文,将其适配到视觉任务具有挑战性。在本文中,我们提出了LiAuto-MindViT,一种新颖的混合视觉骨干网络,它协同了CNN、Mamba和Transformer的优势。我们设计的核心是自适应双向Mamba(ABM),它通过带有可学习alpha混合的双向选择性扫描消除了单向SSM的方向偏差,实现了内容自适应的方向融合,而无需穷举多路径路由的开销。为了进一步加速推理,我们提出了一种部署友好的重参数化ConvSE(RepConvSE)模块,该模块利用结构重参数化来减少延迟和内存访问开销。大量实验表明,LiAuto-MindViT在图像分类、目标检测和语义分割任务上达到了最先进的性能,同时通过重参数化实现了高效推理。

英文摘要

While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local neighborhood correlations and multi-directional spatial contexts for visual understanding. In this paper, we present LiAuto-MindViT, a novel hybrid vision backbone that synergizes the strengths of CNNs, Mamba, and Transformers. The core of our design is the Adaptive Bidirectional Mamba (ABM), which eliminates the directional bias of unidirectional SSMs through bidirectional selective scanning with learnable alpha blending, enabling content-adaptive directional fusion without the overhead of exhaustive multi-path routing. To further accelerate inference, we propose a deployment-friendly Reparameterized ConvSE (RepConvSE) module that leverages structural reparameterization to reduce latency and memory access overhead. Extensive experiments demonstrate that LiAuto-MindViT achieves state-of-the-art performance on image classification, object detection, and semantic segmentation while enabling efficient inference through reparameterization.

发表机构

  • Li Auto Inc.(理想汽车公司)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑