arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09816cs.RO

用于零样本物体目标导航的分层快慢ReAct智能体

Hierarchical Fast-Slow ReAct Agent for Zero-Shot Object-Goal Navigation

Zhaochen Lan, Zhi Yang, Yuxiang Fu, Mengxiang Lin

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对零样本物体目标导航,提出分层快慢ReAct智能体,结合价值图控制器与坐标锚定记忆及VLM,在HM3D、MP3D验证集上取得零样本方法最高成功率,远前沿审议可提升性能。

中文摘要 AI 辅助

零样本物体目标导航(ZSON)要求机器人在从未进入过的建筑中找到指定类别的物体。主流方法使用视觉-语言价值图对前沿区域进行评分:每一步决策都是对当前价值图的argmax,且评分背后的证据在决策时被丢弃。将大型视觉语言模型(VLM)置于感知-动作循环中的系统通常仅根据当前视图按固定频率查询VLM;机器人数分钟前走过的房间不会被重新考虑,且查询失败时没有明确的 fallback(回退)机制。我们将机器人已观测到的内容作为审议对象。我们的分层快慢智能体让价值图控制器每一步都运行,在移动时写入坐标锚定的记忆:包含房间类型和已确认物体实例的语义网格,以及带姿态标签的关键帧的有限存储。VLM在每个候选检测结果写入前进行筛选。审议层在有限的推理-检索-行动循环中读取该记忆,它会在反应层计算出的结构事件触发时唤醒,先基于文本推理,仅当文本无法区分候选时才召回第一人称视图。每次调用和每次运行都有调用次数上限,无调用的第一层可解决最常见的停滞问题,任何失败都会将控制权交还给反应控制器。我们的系统在HM3D v1验证集上达到68.75%的成功率(SR),在MP3D验证集上达到47.29%,是此处比较的零样本方法中最高的成功率。成对比较所有2000个HM3D情节时,通过argmax选择远前沿而非审议会使成功率下降3.40个百分点(95%置信区间[1.70,5.05]);对每个前沿进行审议无法恢复该性能。

英文摘要

Zero-shot object-goal navigation (ZSON) requires a robot to find a named object category in a building it has never entered. The prevailing approach scores frontiers with a vision-language value map: every decision is another argmax over the map as it currently stands, and the evidence behind that score is discarded the moment it is taken. Systems that place a large vision-language model inside the perception-action loop typically query it on a fixed schedule from the current view alone; a room the robot walked through minutes earlier is never reconsidered, and a failed call has no defined fallback. We turn what the robot has already seen into the object of deliberation. Our hierarchical fast-slow agent leaves the value-map controller running at every step and writes a coordinate-anchored memory as it moves: a semantic grid of room types and confirmed object instances, together with a bounded store of pose-tagged keyframes. A VLM screens each candidate detection before it is written. A deliberative layer reads this memory in a bounded reason-retrieve-act loop. It wakes on structural events the reactive layer computes, reasons first over text, and recalls a first-person view only for candidates that text alone cannot separate. Per-invocation and per-run caps bound its calls, a call-free first tier resolves the most frequent stall, and any failure returns control to the reactive controller. Our system reaches 68.75% SR on HM3D v1 val and 47.29% on MP3D val, the highest success rate among the zero-shot methods compared here. Choosing among far frontiers by argmax instead of deliberating costs 3.40 SR points in a paired comparison over all 2000 HM3D episodes (95% CI [1.70, 5.05]); deliberating over every frontier does not recover them.

补充信息

↑