arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

NavCMPO:用于自适应导航的评论家引导的平均流策略优化

NavCMPO: Critic-Guided MeanFlow Policy Optimization for Adaptive Navigation

Junjie An, Yi Wu, Xiao Liu, Yiqun Zhou, Yuechen Wu, Xiaoqing Guan, You Wang, Guang Li

arXiv 2607.14643首次发表:更新:

发表机构

Zhejiang University; Shandong University(浙江大学; 山东大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对无地图视觉导航中策略的问题,提出NavCMPO框架,结合平均流轨迹生成、评论家引导细化和强化学习微调。在InternVLA - N1基准上有高成功率且降低推理延迟,并在Unitree Go2上实现有效模拟到现实的迁移。

AI 中文摘要

基于端到端扩散的策略在无地图视觉导航中表现出强大性能,但迭代去噪过程会带来显著推理延迟,行为克隆则将性能限制在专家演示质量上。我们提出NavCMPO,这是一个两阶段自适应导航框架,结合了几步平均流轨迹生成、评论家引导的细化和强化学习微调。预训练期间,障碍物接近预测任务促使视觉表示捕获障碍物感知空间信息。为弥补几步生成导致的避障能力下降,评论家引导轨迹细化(CGTR)利用由障碍物点云监督训练的评论家的梯度来细化中间轨迹。在适应阶段,平均流策略使用近端策略优化和行为克隆正则化进行微调,同时更新评论家以适应特定实体的观察变化。在InternVLA - N1基准上的匹配训练预算下,NavCMPO实现了74.7%的平均成功率,比重新训练的NavDP基线高出6.4个百分点,同时将推理延迟从85毫秒降至60毫秒。在Unitree Go2上的实验进一步证明了有效的模拟到现实的迁移。

英文摘要

End-to-end diffusion-based policies have demonstrated strong performance in mapless visual navigation, but their iterative denoising process introduces substantial inference latency, while behavior cloning limits performance to the quality of expert demonstrations. We present NavCMPO, a two-stage adaptive navigation framework that combines few-step MeanFlow trajectory generation, critic-guided refinement, and reinforcement learning fine-tuning. During pre-training, an obstacle proximity prediction task encourages the visual representation to capture obstacle-aware spatial information. To compensate for the degradation in obstacle avoidance caused by few-step generation, Critic-Guided Trajectory Refinement (CGTR) uses gradients from a critic trained with obstacle-point-cloud supervision to refine intermediate trajectories. During adaptation, the MeanFlow policy is fine-tuned using Proximal Policy Optimization with behavior-cloning regularization, while the critic is updated to accommodate embodiment-specific observation changes. Under a matched training budget on the InternVLA-N1 benchmark, NavCMPO achieves an average success rate of 74.7\%, exceeding the retrained NavDP baseline by 6.4 percentage points, while reducing inference latency from 85\,ms to 60\,ms. Experiments on a Unitree Go2 further demonstrate effective sim-to-real transfer.

CommentsAccepted for presentation at IROS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑