基于Transformer的令牌融合与动态图规划用于视听导航
Transformer-Based Token Fusion and Dynamic Graph Planning for Audio-Visual Navigation
- Xinjiang University(新疆大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对视听导航中视觉信息缺失导致规划低效的问题,提出TDGP模型,融合Transformer多模态令牌融合与动态图规划,通过碰撞惩罚强制重规划,在Replica和MP3D数据集上超越基线并提升泛化能力。
中文摘要 AI 辅助
视听导航(AVN)要求智能体仅依靠视觉观察和声学线索,定位并导航至一个持续发声的目标。目前,系统在面对不完整或误导性的视觉感知时,缺乏自适应纠正和重新规划的能力。此外,依赖物理碰撞来弥补缺失的视觉信息会导致导航效率低下且不安全,而现有方法过度依赖被动的视觉感知。为解决这些问题,我们提出了基于Transformer的令牌融合与动态图规划(TDGP)模型,该模型融合了高层感知层,并利用Transformer模型融合多模态线索以实现精确的局部规划。接着,设计了一个低层规划层,该层使用物理碰撞惩罚来实时移除与地图发生碰撞的边并施加相应的惩罚,迫使智能体自动重新规划以弥补视觉信息的不足。实验表明,我们的TDGP模型在Replica和Matterport3D(MP3D)数据集上优于基线模型,并且该模型的声增强策略显著提高了在未见声学场景中的泛化能力。
英文摘要
Audio-Visual Navigation (AVN) requires an agent to localize and navigate toward a continuously vocalizing target relying solely on visual observations and acoustic cues. Currently, systems lack the ability to adaptively correct and replan when faced with incomplete or misleading visual perception. Furthermore, relying on physical collisions to compensate for missing visual information results in inefficient and unsafe navigation, whereas existing methods are overly dependent on passive visual perception. To address these issues, we propose the Transformer-based Token Fusion and Dynamic Graph Planning (TDGP) model, which incorporates high-level perception layers and leverages the Transformer model to fuse multimodal cues for precise local planning. Next, a low-level planning layer is designed that uses physical collision penalties to remove edges that collide with the map in real time and apply corresponding penalties, forcing the agent to automatically re-plan to compensate for the lack of visual information. Experiments show that our TDGP model outperforms baseline models on the Replica and Matterport3D (MP3D) datasets, and that the model's sound enhancement strategy significantly improves generalization in unheard acoustic scenarios.