HRO:用于基于大语言模型的零样本目标导航的分层房间到物体框架
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
- School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
- Joint Research Laboratory for Embodied Intelligence, Xinjiang University(新疆大学具身智能联合研究实验室)
- Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(新疆大学丝绸之路多语言认知计算国际联合研究实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对零样本目标导航问题,提出LLM驱动的分层房间到物体(HRO)框架,引导智能体粗到细探索导航至目标物体,实验表明该框架在Gibson和HM3D数据集上优于现有基于LLM的方法。
AI中文摘要:
零样本目标导航旨在让智能体在陌生环境中探索并导航到未知类别的物体,无需特定目标训练。现有基于大语言模型(LLMs)的零样本目标导航方法仅将LLMs用作直接关联物体或区域的平面推理工具,缺乏对类人房间语义到物体定位的分层空间认知建模。本文提出LLM驱动的分层房间到物体(HRO)框架,引导智能体从粗到细探索并导航到目标物体。在Gibson和HM3D数据集上的实验验证了HRO框架优于现有基于LLM的方法,凸显了LLMs在零样本目标导航中的强大潜力。
英文摘要:
Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-trained large models are usually employed to leverage their prior knowledge for guiding the agent's navigation. However, existing zero-shot object-goal navigation methods based on large language models (LLMs) merely utilize LLMs as flat reasoning tools to directly associate objects or regions. They lack the hierarchical spatial cognition modeling of human-like room semantics to object localization, which leads to strong blindness in exploration, insufficient accuracy in semantic association, and failure to fully unleash the common-sense reasoning potential of LLMs. This paper proposes an LLM-driven hierarchical room-to-object (HRO) framework for zero-shot object-goal navigation, which guides the agent to explore and navigate to the target object in a coarse-to-fine manner. Experiments on Gibson and HM3D datasets verify that our HRO framework achieves superior success rate and generalization over existing LLM-based methods, underscoring LLMs' strong potential for zero-shot object-goal navigation.