arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探索性、交流性和可部署性:用于开放世界移动操作的视觉驱动具身智能体

Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation

Boyu Mi, Mengchen Ma, Yifei Yao, Xing Gao, Junting Chen, Yangzi Li, Zihou Zhu, Guohao Li, Zhenfei Yin, Tai Wang, Yao Mu, Jiangmiao Pang, Hanqing Wang

arXiv 2607.13653首次发表:更新:

AI 中文总结

研究针对具身智能体现实部署难题,提出REAL框架,通过建立一致环境API、设计任务组合优化性能,引入REAL - Bench评估。实验显示其智能体在交互式任务中表现出色,在物理机器人上也实现高成功率,弥合现实差距。

AI 中文摘要

具身智能体在现实世界中的部署需要积极探索、视觉基础和交互式意图消歧。然而,现有框架往往依赖特权模拟器状态或假设完整指令,绕过了现实部署挑战。为弥合这一差距,我们提出了REAL,一个用于开放世界移动操作的智能体框架。REAL建立了无预言感知的模拟到现实一致的环境API,并集成了模拟用户以实现人在回路的交互。在这个环境中,我们设计了多样的任务组合来驱动数据收集、监督微调以及在线强化学习,系统地优化智能体性能。为全面评估这种方法,我们引入了REAL - Bench,一个涵盖主动探索、视觉干扰、关节操作和交互式消歧等241个任务的基准。实验结果表明,我们训练的智能体在交互式任务上优于领先的商业闭源语言模型,成功率为56.9%。进一步的实证分析表明,我们的分层训练管道成功地对齐了模型的工具使用能力,同时在扩展的探索视野下保持了强大的开放词汇推理能力。最后,我们在物理双臂移动机器人上部署并评估了我们的框架,在60个真实世界情节中实现了78.3%的端到端成功率。这些物理试验证明了对未见家庭场景的强大零样本可转移性,验证了我们的模拟到现实一致的设计成功地弥合了长期移动操作的现实差距。代码可在指定网址获取。

英文摘要

Real-world deployment of embodied agents requires active exploration, visual grounding, and interactive intent disambiguation. However, existing frameworks often rely on privileged simulator states or assume complete instructions, bypassing realistic deployment challenges. To bridge this gap, we present REAL, an agentic framework for open-world mobile manipulation. REAL establishes sim-to-real-consistent environment APIs without oracle perception and integrates a simulated user to enable human-in-the-loop interaction. Within this environment, we design diverse task compositions to drive data collection, supervised fine-tuning, and online reinforcement learning, systematically optimizing agent performance. To comprehensively evaluate this approach, we introduce REAL-Bench, a benchmark spanning 241 tasks across active exploration, visual distraction, articulated manipulation, and interactive disambiguation. Experimental results demonstrate that our trained agent outperforms leading commercial closed-source VLMs on interactive tasks with a 56.9% success rate. Further empirical analysis reveals that our hierarchical training pipeline successfully aligns the model's tool-use capabilities while maintaining robust open-vocabulary reasoning under extended exploration horizons. Finally, we deploy and evaluate our framework on a physical dual-arm mobile robot, where it achieves a 78.3% end-to-end success rate over 60 real-world episodes. These physical trials demonstrate robust zero-shot transferability to unseen household scenarios, validating that our sim-to-real-consistent design successfully bridges the reality gap for long-horizon mobile manipulation. Code is available at https://github.com/InternRobotics/REAL.

CommentsAccepted to ECCV 2026. 57 pages. Code available at https://github.com/InternRobotics/REAL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑