arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEGNav:面向多模态终身目标导航的时空事件图推理

STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation

Yang Chen, Zhenyu Huang, Wenbo Fu, Danyang Peng, Shi-Yu Tian, Kun-Yang Yu, Lan-Zhe Guo

arXiv 2608.28279首次发表:更新:

发表机构

State Key Laboratory of Novel Software Technology, Nanjing University; School of Intelligence Science and Technology, Nanjing University; School of Artificial Intelligence, Nanjing University(南京大学计算机软件新技术国家重点实验室; 南京大学智能科学与技术学院; 南京大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多模态终身导航方法的局限,本文提出无训练框架STEGNav,通过时空事件图的双轴设计优化导航性能,在多个基准数据集上取得显著效果。

AI 中文摘要

多模态终身导航要求智能体在自主探索未知环境的同时,依次完成由目标对象类别、语言描述或参考图像指定的导航任务。现有方法主要通过构建以状态为中心的语义场景图来完成这些任务,但由于将场景图视为语义观测的持久存储库,这类方法难以区分相似实例、无法联合表示语义目标与探索边界,也无法有效利用导航记忆与轨迹经验。为解决这些局限,本文提出STEGNav(Spatio-Temporal Event Graph Navigation),这是一个无训练框架,沿互补的空间与时间轴将传统场景图扩展为时空事件图。空间轴执行查询条件实例定位,联合表示语义目标与可达性、路径代价及探索效用的占用感知探索边界;时间轴采用感知轨迹的双窗口记忆,保留近期决策-轨迹事件与经验证的跨子任务导航结果。基于VLM的导航智能体对生成的时空事件图进行推理,选择目标实例或探索边界作为下一个导航目标。STEGNav在GOAT-Bench上达到66.3%的成功率(SR)与39.7的路径成功率(SPL),在HM3Dv1和HM3Dv2上分别达到64.0%和69.4%的SR。 ablation研究与误差分析验证了两个轴的互补效应,表明事件驱动的时空表示可提升导航可靠性与跨子任务经验复用。

英文摘要

Multimodal lifelong navigation requires an agent to autonomously explore unseen environments while sequentially completing navigation tasks specified by object categories, language descriptions, or reference images. Existing methods primarily accomplish these tasks by constructing state-centric semantic scene graphs. By treating scene graphs as persistent repositories of semantic observations, these methods struggle to distinguish similar instances, jointly represent semantic targets and exploration frontiers, and effectively exploit navigation memory and trajectory experience. To address these limitations, we propose Spatio-Temporal Event Graph Navigation (STEGNav), a training-free framework that extends conventional scene graphs into spatio-temporal event graphs along complementary spatial and temporal axes. The spatial axis performs query-conditioned instance grounding and jointly represents semantic targets and occupancy-aware exploration frontiers characterized by reachability, path cost, and exploration utility. The temporal axis employs trajectory-aware dual-window memory to retain recent decision--trajectory events and verified cross-subtask navigation outcomes. A VLM-based navigation agent reasons over the resulting spatio-temporal event graph and selects either a target instance or an exploration frontier as its next navigation goal. STEGNav achieves 66.3% SR and 39.7 SPL on GOAT-Bench, as well as SR scores of 64.0% and 69.4% on HM3Dv1 and HM3Dv2, respectively. Ablation studies and error analyses validate the complementary effects of the two axes, demonstrating that event-driven spatio-temporal representations improve navigation reliability and cross-subtask experience reuse.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑