发表机构
Institute for Numerical Simulation; University of Bonn(数值模拟研究所; 波恩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人人群导航中现有方法难以表征多样化短期避障策略的问题,提出PDPO框架,生成短程动作块并引入边界约束,在基准测试中取得更好导航成功率。
AI 中文摘要
机器人人群导航需要在密集、动态且多模态的人机交互下做出安全高效的决策。现有强化学习方法通常在每个时间步输出单一的反应式动作,这限制了其表征多样化短期避障策略的能力。我们提出Planning Diffusion Policy Optimization(PDPO,规划扩散策略优化),这是一种离线转在线的强化学习框架,利用扩散策略为人群导航生成短程动作块。PDPO首先在避障演示上进行预训练,之后通过近端策略优化(PPO)在线微调,将去噪过程视为内部决策过程。执行期间,该策略生成五步动作块并以滚动时域方式应用。此外,我们发现常见人群导航基准中存在一个评估缺陷:若没有显式边界约束,学习到的智能体可能离开有效域并绕过密集人群。为解决此问题,我们引入将违反边界视为碰撞的设置。实验表明,PDPO相较于强基线取得了更高的成功率, ablation(消融)实验证明动作块对修改后的有界基准尤为重要。
英文摘要
Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propose Planning Diffusion Policy Optimization (PDPO), an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation. PDPO is first pretrained on collision-avoidance demonstrations and then fine-tuned online with PPO by treating the denoising process as an internal decision process. During execution, the policy generates a five-step action chunk and applies it in a receding-horizon manner. Furthermore, we observe an evaluation artifact in common crowd-navigation benchmarks: without explicit boundary constraints, learned agents may leave the valid domain and bypass dense crowds. To address this, we introduce a setting in which boundary violations are treated as collisions. Experiments show that PDPO obtains an improved success rate over strong baselines, and ablations demonstrate that action chunks are especially important for the modified bounded benchmark.