AI 中文总结
该研究提出RLMM-Flow框架,结合流策略预训练与隐空间强化学习,在移动操作基准上提升任务成功率等指标,同时保留流的快速推理能力。
AI 中文摘要
移动操作需要生成全身动作块,该动作块需同时满足目标到达、避障、底盘运动学约束、机械臂关节限制以及轨迹平滑性。基于流的生成策略提供了一种从专家演示中学习多模态和时间一致运动先验的高效范式,但仅模仿训练无法将策略质量提升至演示分布之外。我们提出RLMM-Flow,一种基于流的移动操作框架,结合专家流策略预训练与隐空间强化学习后训练。该框架首先学习从专家演示中捕获多模态全身运动先验的流策略;随后冻结预训练的流策略,同时隐空间引导网络将其初始噪声导向高价值动作块。为稳定高维隐空间优化,我们在联合训练隐空间评论者和隐空间演员前预热动作空间评论者,并引入从粗到细的隐空间引导,逐步将控制从共享时间范围的隐空间表示扩展至全维残差表示。在移动操作运动规划基准上的实验表明,RLMM-Flow在保留基于流的快速推理的同时,相比仅模仿的流策略和现有强化学习后训练基线,显著提升了任务成功率、避障能力和轨迹质量。
英文摘要
Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, manipulator joint limits, and trajectory smoothness. Flow-based generative policies provide an efficient paradigm for learning multimodal and temporally consistent motion priors from expert demonstrations, but imitation-only training cannot improve policy quality beyond the demonstration distribution. We propose RLMM-Flow, a flow-based mobile manipulation framework that combines expert flow-policy pretraining with latent-space reinforcement learning post-training. The framework first learns a flow policy that captures a multimodal whole-body motion prior from expert demonstrations. The pretrained flow policy is then frozen, while a latent steering network steers its initial noise toward higher-value action chunks. To stabilize high-dimensional latent optimization, we warm up an action-space critic before jointly training the latent critic and latent actor, and introduce coarse-to-fine latent steering that progressively expands control from a horizon-shared latent representation to a full-dimensional residual representation. Experiments on mobile manipulation motion-planning benchmarks show that RLMM-Flow substantially improves task success, collision avoidance, and trajectory quality over imitation-only flow policies and existing reinforcement learning post-training baselines, while preserving fast flow-based inference.