arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OpenBelief-Nav:面向开放词汇语言引导导航的证据保留型对象记忆

OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation

Dinh Tuan Nguyen, Anh Dao, Phuong Nam Dang, Quan-Dung Pham, Tuyen P. Le, Truong Nguyen, Quan Nguyen

arXiv 2608.13923首次发表:更新:

发表机构

VinMotion, Inc.; University of Southern California(VinMotion公司; 南加利福尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对开放词汇语言引导导航的对象记忆缺陷,提出OpenBelief-Nav模型,保留多源证据并支持灵活读出,在多数据集导航任务中优于基线方法,修正策略可提升目标确认成功率。

AI 中文摘要

开放词汇三维场景图为语言引导导航提供了紧凑的语义记忆,但被映射的对象通常仅通过单一融合特征或确定的语义标签来呈现。这种确定性会从任务时间的交互界面中剔除少数与任务相关的假设。我们提出OpenBelief-Nav,这是一种证据保留型对象记忆,它保留观测级短语、可靠性线索和帧掩码来源,同时维持独立的聚合几何与视觉表示。语义相关的短语被整合为与词汇无关的对象信念,任务特定的读出操作可从该信念执行固定词汇投影或自由形式检索。在5个ScanNet200和8个Replica场景上,全信念投影的mIoU分数分别为0.2742和0.2912,而匹配的早期确定性读出的对应分数为0.2393和0.2701。在78次HM3D-YCB导航试验中,共识和早期确定性检索各实现60/78次成功,信念加权检索为58/78次,DualMap为55/78次。在按10个匹配评估案例组织的20次Unitree G1运行中,与仅执行前1个候选的情况相比,允许最多2次已验证候选尝试的修正策略将目标确认成功率从6/10提升至8/10。代码将在接收后在此URL发布。

英文摘要

Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task-relevant hypotheses from the task-time interface. We present OpenBelief-Nav, an evidence-preserving object memory that retains observation-level phrases, reliability cues, and frame-mask provenance while maintaining separate aggregate geometric and visual representations. Semantically related phrases are consolidated into a vocabulary-independent object belief from which task-specific readouts perform fixed-vocabulary projection or free-form retrieval. On five ScanNet200 and eight Replica scenes, full-belief projection achieves mIoU scores of 0.2742 and 0.2912, compared with 0.2393 and 0.2701 for a matched early-commit readout. Across 78 HM3D-YCB navigation trials, consensus and early-commit retrieval each achieve 60/78 successes, compared with 58/78 for belief-weighted retrieval and 55/78 for DualMap. Across 20 Unitree G1 runs organized as 10 matched evaluation cases, a correction policy permitting at most two verified candidate attempts improves target-confirmation success from 6/10 to 8/10 relative to top-1-only execution. Code will be released upon acceptance at https://openbelief-nav.github.io/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑