面向空中目标导航的双层语义-空间信念映射
Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation
浏览论文内容
中文总结 AI 辅助
提出AeroBelief双层语义-空间信念映射框架,通过直觉层与证据层融合VLM观测,结合保守证据资格和稳定区域引导,在UAV-ON基准上取得最优导航性能。
中文摘要 AI 辅助
空中目标导航(ObjectNav)要求无人驾驶飞行器(UAV)利用机载视觉观测在未知室外环境中定位所描述的目标。视觉语言模型(VLM)能够解释开放式的目标描述和视觉观测,但其帧级输出通常带有噪声、稀疏且空间上具有瞬时性。我们提出AeroBelief,一种双层语义-空间信念映射框架,将瞬时的VLM观测转化为持久的空间引导。该框架将广泛的上下文合理性与目标特定证据相分离:直觉层累积场景级语义线索以支持探索,而证据层保留合格的目标特定观测以用于接近和确认。证据门控融合将两层结合为空间信念热点。我们进一步引入目标条件视觉推理与保守的证据资格判定,以提高空间累积前的观测可靠性。同时,自我中心区域引导将四叉树覆盖转化为以UAV为中心、偏航对齐的方向性提议,并通过时间承诺使其稳定。其区域评分独立于语义信念值,保持探索压力并减少重复的低收益搜索。在UAV-ON基准上的实验表明,AeroBelief在比较方法中实现了最佳的整体成功率(SR)、目标成功率(OSR)和路径长度加权成功率(SPL),分别达到21.61%、35.57%和10.62。这些结果支持了持久语义-空间信念、保守证据资格判定和时间稳定区域引导对空中ObjectNav的有效性。
英文摘要
Aerial Object Goal Navigation (ObjectNav) requires an unmanned aerial vehicle (UAV) to locate a described target in an unknown outdoor environment using onboard visual observations. Vision-language models (VLMs) can interpret open-ended target descriptions and visual observations, but their frame-level outputs are often noisy, sparse, and spatially transient. We propose AeroBelief, a dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance. It separates broad contextual plausibility from target-specific evidence: an intuition layer accumulates scene-level semantic cues for exploration, while an evidence layer preserves qualified target-specific observations for approach and confirmation. Evidence-gated fusion combines the two layers into spatial belief hotspots. We further introduce object-conditioned visual reasoning with conservative evidence qualification to improve observation reliability before spatial accumulation. In parallel, egocentric regional guidance converts quadtree coverage into UAV-centered, yaw-aligned directional proposals and stabilizes them through temporal commitment. Its regional scoring is independent of semantic belief values, maintaining exploration pressure and reducing repeated low-gain search. Experiments on the UAV-ON benchmark show that AeroBelief achieves the best reported overall SR, OSR, and SPL among the compared methods, reaching 21.61%, 35.57%, and 10.62, respectively. These results support the effectiveness of persistent semantic-spatial belief, conservative evidence qualification, and temporally stable regional guidance for aerial ObjectNav.