arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CTAN:用于具身视听导航的周期-时间注意力网络

CTAN: Cycle-Temporal Attention Network for Embodied Audio-Visual Navigation

Teng Liu, Yinfeng Yu

arXiv 2609.17420首次发表:更新:

发表机构

Xinjiang University(新疆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视听导航中多模态融合不足的问题,提出CTAN框架,利用AVRCA周期一致性约束和TCMM时间记忆机制增强跨模态交互,在Replica和Matterport3D上显著提升导航成功率。

AI 中文摘要

视听具身导航通过整合视觉输入和声学信息(如深度观测和双耳音频线索),使机器人能够推断声源的位置。核心挑战在于建立跨异构模态(其特征分布不同)的有效语义交互。然而,现有的特征融合策略往往依赖于简单的多模态聚合,因此无法捕捉潜在的几何和语义关系,导致在复杂环境中信息退化。为克服这些局限,本工作提出了周期-时间注意力网络(CTAN),一个专为主动语义增强融合(而非直接的多模态组合)设计的框架。具体而言,所提出的视听重建交叉注意力(AVRCA)模块采用视觉与声学表示之间的双向周期一致性约束,以强化两种模态的空间语义属性,从而促进更稳健的跨模态交互。此外,我们设计了时间跨模态记忆(TCMM)机制,将实时增强的多模态特征与历史上下文动态整合,减少由听觉盲区导致的性能下降。在Replica和Matterport3D基准上的实验结果表明,所提方法在成功率(SR)、路径长度加权成功率(SPL)和场景导航精度(SNA)方面均优于以往的视听导航方法。

英文摘要

Audio-visual embodied navigation equips robots with the capability to infer the locations of sound sources by integrating visual inputs and acoustic information (e.g., depth observations and binaural audio cues). The core challenge lies in establishing effective semantic interactions across heterogeneous modalities (which exhibit distinct feature distributions). Existing feature fusion strategies, however, often rely on simple multimodal aggregation and therefore fail to capture the underlying geometric and semantic relationships, leading to information degradation in complex environments. To overcome these limitations, this work presents the Cycle-Temporal Attention Network (CTAN), a framework designed for active semantic-enhanced fusion (rather than straightforward multimodal combination). Specifically, the proposed Audio-Visual Reconstruction Cross-Attention (AVRCA) module employs a bidirectional cycle-consistency constraint (between visual and acoustic representations) to reinforce the spatial semantic attributes of both modalities, thereby facilitating more robust cross-modal interaction. Additionally, we design a Temporal Cross-Modal Memory (TCMM) mechanism to dynamically integrate real-time enhanced multimodal features with historical context, reducing performance drops caused by auditory dead zones. Experimental results obtained on the Replica and Matterport3D benchmarks indicate that the proposed approach achieves superior performance over previous audio-visual navigation methods in terms of success rate (SR), success weighted by path length (SPL), and scene navigation accuracy (SNA).

CommentsMain paper (6 pages). Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2026 (IEEE SMC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑