迈向视觉语言导航的双脑最小充分表示
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
浏览论文内容
中文总结 AI 辅助
针对视觉语言导航中存在的问题,提出BrainNav框架,包含逻辑锚定模型等三个组件,通过将语义意图与空间感知对齐,提升智能体在复杂任务中的表现,在实验中相比之前最优方法有改进,为鲁棒的视觉语言导航提供有效基础。
中文摘要 AI 辅助
连续环境中的视觉与语言导航(VLN-CE)要求智能体将语言与自我中心观察相结合,并在未见场景中进行规划。尽管近期多模态大模型和基于世界模型的方法改进了导航,但常保留过多无关细节,削弱泛化能力并增加计算负担。我们提出基于最小充分性原则的导航框架BrainNav。它由逻辑锚定模型、极简约束对齐模块和压缩世界模型组成,能将语义意图与空间感知对齐,增强智能体在复杂任务中的鲁棒性和效率。实验表明,BrainNav在R2R-CE和RxR-CE验证集上相比之前的最优方法有提升。这些结果表明最小充分世界表示为鲁棒的VLN提供了有效基础。
英文摘要
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening generalization and increasing computational burden. We propose BrainNav, a navigation framework grounded in the Principle of Minimal Sufficiency. BrainNav consists of three components: a Logical Anchor Model that implements instruction-aware selective perception to suppress environmental noise, a Minimalist Constraint Alignment module that serves as a compact cross-modal bottleneck, efficiently synchronizing discrete linguistic intent with continuous latent dynamics while filtering out redundant information, and a Compression World Model that predicts action-conditioned states within a condensed, low-rank latent space. These modules align semantic intent with spatial perception, enhancing the agent's robustness and efficiency in complex tasks. Experiments show that BrainNav improves over prior SOTA by 2.0 % / 1.0 in SR/SPL on R2R-CE val-unseen and 0.94 % / 0.78 on RxR-CE val-unseen. These results indicate that minimally sufficient world representations provide an effective foundation for robust VLN.
发表机构
- Hangzhou Dianzi University(杭州电子科技大学)
- University of Zurich(苏黎世大学)
- University of Exeter(埃克塞特大学)
机构由 AI 辅助整理,请以论文原文为准。