arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39207cs.RO

ASENA:用于具身导航的自进化智能体

ASENA: Self-evolving Agents for Embodied Navigation

An-Chieh Cheng, Isabella Liu, Edmund Bu, Johan Bjorck, Hongxu Yin, Zhengyi Luo, Jan Kautz, Linxi "Jim" Fan, Yuke Zhu, Sifei Liu

首次发表
浏览论文内容

中文总结 AI 辅助

提出ASENA具身导航系统,将编程智能体与机器人结合,通过自进化工作区提升导航性能,在R2R和RxR基准上达到最先进水平,并支持真实世界任务。

中文摘要 AI 辅助

我们提出了ASENA,一个具身智能体系统,它将通用编程智能体与机器人感知、计算、受监督执行和持久经验连接起来。智能体可以编写和执行程序、检查记录的结果、修复故障,并复用笔记和可执行技能,同时保持其模型权重固定。我们进一步引入了ASENA-VLN,一个4B单目导航策略,作为该可编程系统内的可选工具。ASENA-VLN使用一个共享的视觉-语言解码器,预测扩展路线和短视界行为的本体坐标系轨迹,该解码器在路线指令、视觉问答以及一个新整理的基于几何的原子导航任务数据集上进行训练。作为独立策略,ASENA-VLN在R2R上达到了68.7%的最先进成功率,在RxR上达到了70.2%。当与编程智能体集成时,学习到的导航在两个智能体基准上将ASENA的成功率提高了11个百分点,同时减少了执行时间。通过持久工作区演化和模拟器反馈,在重复的100任务子集上进行的十次遍历进一步将成功率从R2R上的72%提高到98%,从RxR上的65%提高到89%。在具身问答方面,ASENA以更少的交互步骤达到了最先进的准确率。最后,在Unitree G1上的真实世界演示结合了搜索、视觉检查、空间推理和合成手势,无需预先构建地图,说明了在线编程如何将机器人行为扩展到路线跟随和预定义技能之外。

英文摘要

We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weights fixed. We further introduce ASENA-VLN, a 4B monocular navigation policy that serves as an optional tool within this programmable system. ASENA-VLN predicts body-frame trajectories for both extended routes and short-horizon behaviors using a shared vision-language decoder trained on route instructions, visual question answering, and a newly curated dataset of geometry-derived atomic navigation tasks. As a standalone policy, ASENA-VLN achieves state-of-the-art success rates of 68.7% on R2R and 70.2% on RxR. When integrated with a coding agent, learned navigation improves ASENA's success rate by 11 percentage points on both agentic benchmarks while reducing execution time. Through persistent workspace evolution and simulator feedback, ten passes over recurring 100-task subsets further improve success from 72% to 98% on R2R and from 65% to 89% on RxR. On embodied question answering, ASENA achieves state-of-the-art accuracy with fewer interaction steps. Finally, real-world demonstrations on a Unitree G1 combine search, visual inspection, spatial reasoning, and synthesized gestures without a pre-built map, illustrating how online programming extends robot behavior beyond route following and predefined skills.

发表机构

  • NVIDIA(英伟达)
  • University of California, San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑