发表机构
Norwegian University of Science and Technology (NTNU)(挪威科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对具身问答中探索仅依赖局部观察而忽视环境结构先验的问题,提出HFLEX-EQA分层框架,结合场景图、VLM规划、语义前沿与平面图先验,在OpenEQA和ExploreEQA基准及真实机器人上验证了有效性。
AI 中文摘要
具身问答(EQA)要求智能体探索一个先前未见过的环境,收集相关信息,并回答关于该场景的问题。近期的方法利用视觉-语言模型(VLM)以及语义地图或场景图来引导探索。然而,探索通常仅由局部观察驱动,而关于环境的结构性先验信息在很大程度上未被利用。我们提出了HFLEX-EQA,一个分层的EQA框架,该框架结合了在线场景图构建、基于VLM的规划、语义前沿探索和平面图先验。该系统从RGB-D观察中增量构建分层场景图和开放词汇占用地图,使VLM能够联合推理场景图、任务相关的视觉观察、探索历史以及估计的拓扑平面图。此外,我们引入了一种房间发现策略,利用平面图和开放词汇前沿语义,将探索引导至语义相关但目前尚未观察到的房间类型。我们在OpenEQA和ExploreEQA基准上评估了HFLEX-EQA,并展示了在真实室内环境中四足机器人上的部署。我们的结果证明了将基于VLM的分层规划与结构性平面图先验相结合对EQA任务的好处。
英文摘要
Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM- based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.