AECNav:用于高效零样本开放词汇对象导航的主动证据整合
AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation
浏览论文内容
中文总结 AI 辅助
AECNav是一种无需训练的零样本开放词汇对象导航方法,通过证据门控感知、证据整合与主动证据获取三个组件,在多个基准数据集上达到最新性能,且在物理机器人上表现优异。
中文摘要 AI 辅助
开放词汇场景下的零样本对象目标导航(ZSON)极具挑战性,因为它要求机器人在未见过的环境中定位任意指定的对象,且无需针对该任务进行专门训练。当前该任务仍存在延迟高、精度有限的问题,原因在于感知流程存在冗余,且可靠目标确认的证据不足。在本文中,我们将ZSON重新定义为证据驱动的感知到决策问题,并提出AECNav,这是一个无需训练的流水线,由三个组件构成:i)证据门控感知,它在所有推理阶段利用共享编码建立统一的语义基础,消除冗余计算;ii)证据整合,它将检测结果聚合成聚类级对数几率信念,明确区分真实目标支持与视觉相似干扰项的虚假置信度,同时将预期检测的缺失视为负面证据;iii)主动证据获取,它通过选择以最小遍历成本最大化信息增益的前沿区域,在弱语义线索下维持高效探索。结果表明,AECNav显著优于先前方法,在HM3D-v2、HM3D-OVON和MP3D上分别达到84.7%、57.3%和51.3%的最新成功率,同时推理开销大幅降低,在四足物理机器人上以约5Hz运行时,40次试验的成功率达95%。代码将在录用后公开。
英文摘要
Zero-shot object-goal navigation (ZSON) in open-vocabulary scenarios is challenging, as it requires a robot to locate an arbitrarily specified object in an unseen environment without task-specific training. Currently, the task still suffers from high latency and limited accuracy due to redundant perception pipelines and insufficient evidence for reliable target confirmation. In this letter, we reframe ZSON as an evidence-driven perception-to-decision problem and present AECNav, a training-free pipeline built on three components: i) Evidence-gated perception, which utilizes a shared encoding across all reasoning stages to establish a unified semantic basis and eliminate redundant computations; ii) Evidence consolidation, which aggregates detections into cluster-level log-odds beliefs. This explicitly separates genuine target support from the false confidence of visually similar distractors, while treating the absence of expected detections as negative evidence; and iii) Active evidence acquisition, which sustains productive exploration under weak semantic cues by selecting frontiers that maximize information gain at minimal traversal cost. As a result, AECNav significantly outperforms previous methods and achieves state-of-the-art success rates of 84.7%, 57.3%, and 51.3% on HM3D-v2, HM3D-OVON, and MP3D, respectively, with substantially lower inference overhead, and attains 95% success across 40 trials on a physical quadruped robot at roughly 5Hz. Code will be made publicly available upon acceptance.