发表机构
Kyushu University(九州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出DSRL方法,通过固定扩散策略仅训练噪声策略,并在多种子集成基础上实现社交导航中的性能保持在线自适应,验证了高效学习与灵活行为控制。
AI 中文摘要
在社交导航中,建模人类与机器人之间的复杂交互是困难的,因此深度强化学习已被积极研究。然而,由于仅靠仿真无法完全复现多样化的场景、机器人动力学以及随部署环境变化的社会规范,在部署环境中进行微调是有前景的。在此过程中,需要保持基础模型性能的学习,以免损害导航的主要目标,即避开行人并到达目的地。在本研究中,我们提出了一种通过强化学习应用扩散引导(DSRL)的方法,该方法仅训练噪声策略,同时保持扩散策略固定,从而实现保持性能的学习。此外,我们整合了使用多种子训练的基于扩散的强化学习策略以构建基础策略,提高了学习性能。我们的评估表明,与其他方法相比,所提出的方法能够在保持性能的同时实现高效学习,并且我们通过硬件在环仿真确认了通过适应社会规范实现的灵活行为控制,以及其在物理机器人上的有效性。
英文摘要
In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.