arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于视听导航的混合曼巴

A Hybrid Mamba for Audio-Visual Navigation

Yi Wang, Yinfeng Yu

arXiv 2607.13110首次发表:更新:

发表机构

School of Computer Science and Technology, Xinjiang University; Joint Research Laboratory for Embodied Intelligence, Xinjiang University; Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(新疆大学计算机科学与技术学院; 新疆大学具身智能联合研究实验室; 新疆大学丝绸之路多语言认知计算国际联合研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视听导航骨干网络多年未变的问题,提出Samba,用曼巴状态编码器取代传统GRU,构建音频曼巴编码器,实验表明其泛化性能出色,能提升导航成功率,以低成本解锁更强能力,为范式演进提供途径。

AI 中文摘要

自2020年以卷积神经网络和循环架构为中心的范式确立以来,视听导航的基础骨干网络五年多来没有本质变化,不足以支持动态多模态序列的高效表示。本文提出Samba(用于视听导航的混合曼巴)。它使用具有自适应选择功能的曼巴状态编码器(M-SE)取代传统GRU进行时间聚合,并构建音频曼巴编码器(AME)来弥补卷积算子在捕捉频谱图中全局时频依赖性方面的局限性。实验表明,Samba在面对未听过的声源和未见场景时具有出色的泛化性能。在Matterport3D数据集上,与现有最先进模型相比,它将导航成功率(SR)提高了11.3%,在具有更精细场景结构的Replica数据集上性能提升更显著。这种现代化架构重建以更低计算成本解锁更强的具身表示能力,为视听导航领域的范式演进提供了高度稳健的技术途径。

英文摘要

Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks for audio-visual navigation have undergone no essential changes for more than five years, making them inadequate to support efficient representation of dynamic multimodal sequences. This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation). It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (AME) to remedy the limitations of convolutional operators in capturing global time-frequency dependencies in spectrograms. Experiments demonstrate that Samba exhibits exceptional generalization performance when facing unheard sound sources and unseen scenes. On the Matterport3D dataset, it improves the navigation success rate (SR) by 11.3\% compared with existing state-of-the-art models, and the performance gain is even more pronounced on the Replica dataset, which features finer scene structures. Such modernized architectural reconstruction unlocks stronger embodied representation capabilities at a lower computational cost, thereby providing a highly robust technical pathway for paradigm evolution in the field of audio-visual navigation.

CommentsMain paper (6 pages). Accepted for publication by IEEE International Conference on Systems and Man and Cybernetics 2026 (IEEE SMC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑