arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具身智能体自主掌控:最小接口零样本智能体在视觉语言导航中可媲美工业级策略

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

Jian Zhou, Xunyi Zhao, Gengze Zhou, Zerui Li, Sihao Lin, Jiajun Liu, Qi Wu

arXiv 2607.26148首次发表:更新:

发表机构

Australian Institute for Machine Learning; Adelaide University; Responsible AI Research Centre; CSIRO Data61; The University of Queensland(澳大利亚机器学习研究所; 阿德莱德大学; 负责任人工智能研究中心; 联邦科学与工业研究组织数据61中心; 昆士兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出智能具身控制,以零样本导航为测试平台,评估三种智能体框架,发现最小接口下部分框架可媲美工业级策略,同时揭示模型、框架与接口对具身智能体的互补作用。

AI 中文摘要

自主具身智能体必须维持长期决策循环,该循环涉及多步骤的感知、行动、验证与自我修正。当前系统通过特定任务工作流或具身策略维持该循环,本研究探讨第三种形式——智能具身控制,即通用智能体自身掌控该循环。以零样本导航为受控测试平台,在仅配备单目RGB相机与离散动作的严格最小条件下,评估三种软件工程智能体框架:opus-5在默认配置下的平均成功率为70.7±3.5%(三次运行均值),fable-5在最大努力配置下成功率达78%。当将训练好的路径点工具作为可选能力与基础动作一同提供时,混合fable-5智能体在默认配置下成功率达76.7±0.6%,仅需一半环境步骤,且耗时不足最大努力基础动作运行的四分之一。受控干预实验表明,能力效果主要以模型为中心:模型选择对成功率影响显著,框架效果仅具描述性,强制路径点接口可辅助较弱模型但可能阻碍较强模型。不过,该方法在长视野任务上性能骤降,延迟与上下文增长限制了持续运行。这些结果显示,智能控制在零样本导航中已具备竞争力,模型、框架与接口为自主具身智能体提供了互补发展路径。

英文摘要

Autonomous embodied agents must sustain a long decision-making loop that involves perceiving, acting, verifying, and self-correcting over many steps. Current systems sustain this loop through task-specific workflows or embodied policies. However, these fixed workflows and policies offer limited flexibility across environments and often lack effective recovery strategies when execution goes wrong. We find that a general-purpose agent can instead sustain the loop on its own. We term this organization agentic embodied control: the reasoning model directly steers every action, keeping reasoning and control aligned. Using zero-shot navigation as a controlled testbed, we equip three coding-agent harnesses with only a monocular RGB camera and discrete actions. At default effort, replicated opus-5 runs average $70.7\pm3.5$% success, while fable-5 reaches 78% at maximum effort. When a trained waypoint tool is offered alongside primitives, the hybrid fable-5 agent reaches $76.7\pm0.6$% at default effort, using half the environment steps and under a quarter of the wall time. Across the ablations, model choice dominates performance variation. Observed harness differences are modest, and forced waypoints help weaker models but can hinder stronger ones. Although longer horizons, latency, and context growth remain barriers to sustained autonomy, these results show that a general-purpose model can already achieve competitive embodied control without a navigation policy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑