arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MobiAgent:面向长时程移动操作的双循环递归策略自改进

MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

Chenzhi Liu, Yue Zhang, Jiehong Lin, Jianan Wang, Bo Wang, Zhongrui Wang, Xiaojuan Qi

arXiv 2610.03476首次发表:更新:

发表机构

The University of Hong Kong; Astribot; Southern University of Science and Technology(香港大学; Astribot; 南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MobiAgent双循环智能体框架,通过内循环原子技能组合与外循环自动数据回收,解决长时程移动操作中的误差累积和策略自改进问题,在多个基准上显著提升成功率。

AI 中文摘要

长时程移动操作因执行误差累积以及移动与手臂控制之间的能力干扰而面临重大挑战。尽管最近的视觉-语言-动作模型在短时程任务中表现出色,但它们缺乏多阶段目标所需的分层推理能力。此外,现有的分层智能体存在子任务映射僵化、重规划不灵活以及缺乏持续学习等问题。为解决这些局限,我们提出了MobiAgent,一种双循环智能体框架,将鲁棒的部署执行与递归策略自改进相结合。在部署期间,内循环通过高度可组合的原子技能将高层推理与低层控制解耦。它利用视觉-语言模型进行滚动时域规划和视觉反思,动态组合技能以确保鲁棒的误差恢复。这些技能由共享统一VLM骨干的专用流匹配专家执行,在缓解能力干扰的同时最大化可重用性。同时,外循环通过自主分割和验证部署轨迹、聚类以发现原子技能,并持续微调技能库(无需人工标注),驱动自动化终身学习。在RoboCasa、BEHAVIOR-1K和真实世界任务上的评估证明了MobiAgent的有效性。它在BEHAVIOR-1K上比π0.5-TA高出22.5个百分点,并能从执行失败中鲁棒恢复。通过自主数据回收,在RoboCasa上成功率从7.50%提升至27.50%,在Astribot S1上从32.5%提升至57.5%。

英文摘要

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms $π_{0.5}$-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.

CommentsAccepted at the Conference on Robot Learning (CoRL) 2026. Project page: https://kaiknower.github.io/mobiagent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑