MobileVLA-R1 2.0:面向移动机器人控制的强化学习增强推理
MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
浏览论文内容
中文总结 AI 辅助
提出MobileVLA-R1 2.0,通过监督CoT对齐与强化学习增强具身推理,并引入推理条件动作解码器,在VLN-CE、QUARD及真实机器人上显著提升移动操作性能。
中文摘要 AI 辅助
将自然语言指令落地为可靠且可执行的动作,对于移动机器人上的视觉-语言-动作(VLA)系统而言仍是一项基本挑战,其原因在于高层语义推理与底层运动及操作控制之间持续存在的鸿沟。现有方法往往依赖隐式推理或整体式动作预测,这使得在生成精确且适应性强的机器人动作的同时,维持连贯的长时程决策变得困难。为应对这一挑战,我们提出了MobileVLA-R1 2.0,一个强化学习增强的VLA框架,它明确地将结构化具身推理与可执行的移动机器人控制耦合起来。该框架通过监督式思维链(CoT)对齐和强化学习,学习对具身轨迹进行多粒度推理,从而在纯行为监督之外提升推理到动作的一致性。为同时支持移动和操作,我们进一步引入了一个推理条件动作解码器,将多模态推理表示映射到任务级动作目标,随后由机器人控制器将其转换为具身特定的指令。这一设计提供了统一的感知-推理-动作接口,同时将高层动作生成与机器人特定的驱动解耦。我们在语言引导导航、四足控制和仿人移动操作上进行了广泛评估,涵盖VLN-CE、QUARD以及Unitree Go2和G1机器人的真实世界部署。MobileVLA-R1 2.0持续优于强VLA基线,在VLN-CE上平均SR提升1.6个百分点,在真实世界G1移动操作任务上相比MobileVLA-R1全任务成功率提升10.0个百分点,并在不同机器人平台上展示了稳健的长时程指令跟随和闭环执行能力。
英文摘要
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semantic reasoning and low-level locomotion and manipulation control. Existing approaches often rely on implicit reasoning or monolithic action prediction, making it difficult to maintain coherent long-horizon decision making while producing precise and adaptable robot actions. To address this challenge, we propose MobileVLA-R1 2.0, an RL-enhanced VLA framework that explicitly couples structured embodied reasoning with executable mobile robot control. The framework learns multi-granularity reasoning over embodied trajectories through supervised Chain-of-Thought (CoT) alignment and reinforcement learning, improving reasoning-to-action consistency beyond purely behavioral supervision. To support both locomotion and manipulation, we further introduce a reasoning-conditioned action decoder that maps multimodal reasoning representations to task-level action targets, which are subsequently translated into embodiment-specific commands by robot controllers. This design provides a unified perception-reasoning-action interface while decoupling high-level action generation from robot-specific actuation. We conduct extensive evaluations on language-guided navigation, quadruped control, and humanoid mobile manipulation, covering VLN-CE, QUARD, and real-world deployments on Unitree Go2 and G1 robots. MobileVLA-R1 2.0 consistently outperforms strong VLA baselines, achieving an average 1.6 point improvement in SR on VLN-CE and a 10.0 point improvement in full-task success on real-world G1 mobile manipulation tasks over MobileVLA-R1, while demonstrating robust long-horizon instruction following and closed-loop execution across different robotic platforms.
发表机构
- Peking University(北京大学)
- South China University of Technology(华南理工大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。