GPT-6-Astra 点亮具身导航:连续环境中零样本视觉-语言导航评估
GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
查看机构详情
- Singapore Management University(新加坡管理大学)
- Australia Institute for Machine Learning(澳大利亚机器学习研究所)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究评估 GPT-6-Astra 在连续环境中使用单目 RGB 进行零样本视觉-语言导航,无需微调或地图,在 R2R-CE 基准上达到 79.0% 成功率,超越现有方法,并指出路线执行与目标验证仍是挑战。
中文摘要 AI 辅助
我们研究 GPT-6-Astra(一种通用基础模型)是否能够利用自身的感知、推理和决策能力在陌生环境中导航。我们的评估聚焦于通过 Codex 环境中的最小接口在连续环境中进行零样本视觉-语言导航(VLN-CE),旨在释放 GPT-6-Astra 在导航方面的全部潜力。使用单目 RGB 图像,GPT-6-Astra 决定何时观察、如何移动以及何时停止,无需导航特定的微调、训练好的航路点预测器或预构建的场景地图。我们的评估得出四个主要发现。第一,GPT-6-Astra 仅使用单目 RGB 观测即可实现强大的零样本导航性能。在广泛采用的零样本 R2R-CE 基准上,超推理实现了 79.0% 的成功率,分别超过最强的已报告零样本和受监督成功率 13.0 和 6.9 个百分点。第二,GPT-6-Astra 通过建立空间关系、跟踪任务进度和修正其行动,将多阶段语言指令推进为连贯、自适应的导航。第三,即使采用超推理,可靠的路线执行和目标验证仍然具有挑战性。看似合理的局部地标匹配并不总能导致正确的任务完成。第四,这些能力促使我们重新思考具身学习的作用。未来的 VLN 研究应基于基础模型,以推进通用且可靠的具身智能。
英文摘要
We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focuses on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra's full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four key findings and implications. First, GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations. On R2R-CE-100, GPT-6-Astra (ultra reasoning) achieves a success rate of 81.3%, exceeding the strongest reported zero-shot and even train-based success rates by 15.3 and 9.2 percentage points, respectively. Second, GPT-6-Astra exhibits promising capabilities in interpreting multi-stage instructions, understanding the environment, and adjusting routes. Third, execution and goal-verification failures persist even with ultra reasoning. Fourth, these results motivate combining general-purpose model capabilities with navigation-specific expertise. Based on these findings, future VLN research should investigate which aspects of instruction interpretation, spatial understanding, and navigation decision-making general-purpose models can handle directly, and where navigation-specific learning can extend their capabilities. This includes exploring how spatial representations, navigation experience, and learned control skills can improve progress tracking, error recovery, and goal verification while preserving the flexibility to adjust routes.