Beyond Textual Knowledge-Leveraging Multimodal Knowledge Bases for Enhancing Vision-and-Language Navigation
超越文本知识的多模态知识库用于增强视觉-语言导航
机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) ; Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语言认知计算联合国际研究实验室)
专题命中 GUI与网页智能体 :agent(abstract);分类 cs.AI
AI总结 本文提出BTK框架,通过整合环境特定文本知识与生成图像知识库,提升视觉-语言导航中的语义 grounding 和跨模态对齐,实验表明在R2R和REVERIE数据集上性能显著提升。
Comments Main paper (37 pages). Accepted for publication by the Information Processing and Management,Volume 63,Issue 6,September 2026,104766