arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAIN:面向移动机器人的、结合主动对话 grounding 的结构感知交互式导航

SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

Yuhao Cao, Xiao Liu, Yang Xie, Lu Liu, Haoyao Chen

arXiv 2608.09196首次发表:更新:

发表机构

School of Mechanical Engineering and Automation, Harbin Institute of Technology, Shenzhen; Department of Mechanical Engineering, City University of Hong Kong(哈尔滨工业大学(深圳)机械工程与自动化学院; 香港城市大学机械工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SAIN零样本框架,将主动对话转化为持久导航状态,在VL-LN IIGN基准上提升了导航成功率与加权成功率,无需特定任务策略训练,验证了对话转状态机制的有效性。

AI 中文摘要

现有多数视觉-语言导航任务假设指令完整且无歧义,但现实中机器人常遇到模糊、欠规范或不完整的自然人类指令,需通过主动提问解决此类不确定性。交互式实例目标导航(Interactive Instance Goal Navigation, IIGN)要求具身智能体通过主动对话,在模糊的类别级指令下找到特定实例。然而,现有对话驱动方法常将神谕答案作为即时决策的瞬态文本上下文,而非持久的空间或以物体为中心的结构化状态。本文提出SAIN,这是一种零样本框架,可将主动对话转化为持久的导航状态。SAIN并非将神谕答案作为单步文本提示使用,而是将其编译为目标证据、路线级走廊记忆和物体候选标签;这些状态被存储在结构化的价值记忆、房间记忆、图记忆和物体记忆中,随后由统一策略用于前沿排序和最终目标接近。在VL-LN IIGN基准上,SAIN相较于已报告的最强对话驱动基线,将成功率(SR)从20.2提升至25.4,将加权成功率(SPL)从13.07提升至14.17,且无需进行特定任务的策略训练。结果表明,对话到状态的转换是长程交互式实例导航的有效零样本机制。项目网站:this https URL

英文摘要

Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural human instructions that are ambiguous, underspecified, or incomplete, requiring them to resolve such uncertainties through active questioning. Interactive Instance Goal Navigation (IIGN) requires an embodied agent to find the specific instance under an ambiguous category-level instruction through active dialogue. However, existing dialogue-enabled methods often consume oracle answers as transient textual context for immediate decisions, rather than persistent spatial or object-centric structured state. We present SAIN, a zero-shot framework that turns active dialogue into persistent navigation state. Instead of consuming oracle answers as one-step text hints, SAIN compiles them into target evidence, route-level corridor memory, and object-candidate labels. These states are stored in structured value, room, graph, and object memories, then consumed by a unified policy for frontier ranking and final target approach. On the VL-LN IIGN benchmark, SAIN improves SR from 20.2 to 25.4 and SPL from 13.07 to 14.17 over the strongest reported dialogue-enabled baseline, while requiring no task-specific policy training. The results support dialogue-to-state conversion as an effective zero-shot mechanism for long-horizon interactive instance navigation. Project website: https://zorattc.github.io/SAIN/

CommentsThe authors have decided to withdraw this preprint due to unresolved internal disagreements regarding its public release at this stage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑