LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
LaViRA: 语言-视觉-机器人动作翻译用于连续环境中的零样本视觉语言导航
机构 * School of Computer Science, Nanjing University(南京大学计算机科学学院) ; School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 GUI与屏幕智能体 :grounding(abstract);multimodal large language model(abstract)
AI总结 LaViRA通过分解动作层次结构,利用多模态大语言模型的优势,实现连续环境中零样本视觉语言导航的高效导航与泛化能力。
Comments ICRA 2026