GPT-6-Astra 在导航工作流中的应用:连续环境中零样本视觉与语言导航的行为分析
How Far Can GPT-6-Astra Go? Evaluating Capabilities in Zero-Shot Vision-and-Language Navigation
浏览论文内容
中文总结 AI 辅助
本研究分析 GPT-6-Astra 在零样本 VLN-CE 导航系统中的行为,发现其成功率 52.0%,但存在局部判断与自主完成之间的差距,核心挑战在于将正确判断转化为持续进展和适当停止。
中文摘要 AI 辅助
我们研究了 GPT-6-Astra 在连续环境中的零样本视觉与语言导航(VLN-CE)系统中的应用,该系统负责解释指令、评估周围环境并提出行动建议。该系统采用常见的观察-决策-执行工作流,直接调用模型 API,无需封装智能体框架或针对导航的微调。在此工作流中,每次请求都会收到选定的观察结果、执行反馈和保留的进度记录。评估覆盖整个系统,包括上下文管理和行动控制。我们在 Open-Nav 使用的 100 个 R2R-CE 验证未见过的片段中选取了 50 个进行评估。该系统达到了 52.0% 的成功率、48.9% 的 SPL 和 70.8% 的 nDTW。我们的分析突出了三个发现。第一,记录的回答通过观察和提供的历史信息将地标和早期行动与指令联系起来。第二,审查包括对额外视图的请求和对不确定判断的修正。第三,结果表明任务理解与自主完成之间存在差距:在旋转继续时,未完成的穿越被识别。在终止时,36.0% 的片段通过工作流接受的停止(STOP)成功,另有 16.0% 在步数限制时满足距离标准。这些结果突出了一个核心挑战:将正确的局部判断转化为持续的进展和适当的停止。
英文摘要
We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation--decision--execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control. We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0\%, an SPL of 48.9\%, and an nDTW of 70.8\%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues. At termination, 36.0\% of episodes succeed with a workflow-accepted STOP, while another 16.0\% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.
发表机构
- Singapore Management University(新加坡管理大学)
- Australia Institute for Machine Learning(澳大利亚机器学习研究所)
机构由 AI 辅助整理,请以论文原文为准。