发表机构
Fudan University; Shanghai AI Laboratory; Sun Yat-sen University; Tsinghua University(复旦大学; 上海人工智能实验室; 中山大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对导航扩散策略泛化性不足的问题,提出GQRM框架对X-NavDP进行后训练,大幅提升了跨实体视觉导航的模拟与现实场景成功率。
AI 中文摘要
导航扩散策略的预训练依赖大规模专家演示数据,这些数据通常由适配单一标称机器人的全知规划器生成,限制了策略对多样实体及需多样局部反应行为的具挑战性场景(如摆脱死胡同或绕过长障碍物,仅利用机载局部观测)的泛化能力。用强化学习(RL)对策略进行后训练是合理解决方案,但此前针对扩散模型的RL方法仅能带来微小改进,原因在于扩散策略的难处理似然性会导致策略梯度不稳定,且策略探索效率低下。为解决这些挑战,我们提出数据高效的扩散RL后训练框架GQRM(Group Q-score Reweighted Matching,组Q分数重加权匹配),该框架引入两种互补设计:(i)带行为扰动的自举探索策略,可保留预训练策略的先验;(ii)组Q分数归一化机制,对每个状态计算每条轨迹的值以实现高效重加权分数匹配。通过在异构实体上进行分布式在线RL训练,得到的微调策略X-NavDP实现了最先进的跨实体视觉导航性能,在模拟环境中整体成功率从61.20%提升至84.28%,在现实世界困难场景中从10%提升至65%。代码和模型已公开于此httpsURL。
英文摘要
Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suited to a single nominal robot. This limits the policy's generalization to diverse embodiments and challenging scenarios (e.g., escaping dead ends or detouring long obstacles) that demand diverse local reactive behaviors with only onboard local observations. Post-training the policy with reinforcement learning (RL) offers a principled remedy. However, previous RL for diffusion approaches lead to only marginal improvements. This is because the intractable likelihood of diffusion policies renders policy gradients unstable in addition to inefficient policy exploration. To address these challenges, we propose a data-efficient diffusion RL post-training framework - GQRM (Group Q-score Reweighted Matching). Our framework introduces two complementary designs: (i) a self-bootstrapped exploration strategy with behavior perturbation that preserves the pretrained policy prior, and (ii) a group Q-score normalization mechanism that computes per-trajectory values on each state for efficient reweighted score matching. By conducting distributed online RL training across heterogeneous embodiments, the resulting fine-tuned policy, X-NavDP, achieves state-of-the-art cross-embodiment visual navigation performance, improving the overall success rate from 61.20% to 84.28% in simulation and 10% to 65% in real-world hard cases. The code and model are publicly available at https://yty-sky.github.io/x-navdp-project-page.
Comments20 pages, 4 figures