arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向社交导航的优势驱动型显式记忆

Advantage-Driven Explicit Memory for Social Navigation

Yeonsoo Park, Mattia Racca, Guillaume Bono, Steeven Janny, Gianluca Monaci, Tomi Silander, Christian Wolf

arXiv 2608.25610首次发表:更新:

发表机构

Seoul National University; Naver Labs Europe(首尔大学; Naver欧洲实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对社交导航任务,提出将非参数化显式记忆集成到循环PPO架构,利用RL优势信号保留关键事件,提升智能体泛化能力与对分布外社交行为的鲁棒性。

AI 中文摘要

机器人策略主要通过经典的参数化模仿学习或强化学习(RL)变体进行训练,训练过程中智能体的行为仅存储在策略的网络参数中,这给表示学习算法带来了沉重负担。我们提出一种配备非参数化记忆的新型导航智能体,该记忆可显式索引导致关键事件的先前步骤,其优势有两点:一是允许策略将部分行为外包至显式记忆中;二是通过让智能体在部署期间从测试回合收集数据,鼓励一种持续学习的形式,从而更好地泛化到分布外(OOD)场景。在社交导航场景中,我们表明这提升了智能体保留稀疏、高代价失败案例(如与人类碰撞)的能力。若策略在仿真环境中训练,还能通过将部分决策建立在真实数据基础上,部分弥合仿真到真实的差距。我们将显式记忆集成到循环近端策略优化(recurrent PPO)架构中,并利用隐藏状态进行记忆检索以捕捉连续时空动态。通过利用RL智能体的优势信号,实现了利用稀有高影响事件的目标。我们在仿真环境中训练智能体,采用光真实感渲染与非视觉人群模拟相结合的方式,结果表明该智能体对OOD社交行为具有鲁棒性。

英文摘要

Robot policies are predominantly learned with classical parametric variants of imitation learning or RL, where training stores the agent's behavior exclusively in the policy's network parameters, putting a heavy burden on the representation learning algorithm. We propose a new navigation agent equipped with non-parametric memory which explicitly indexes prior steps leading to critical events. The advantages are twofold: first, it allows the policy to outsource some of its behavior into an explicit memory; second, it encourages a form of continual learning by allowing an agent to collect data from its testing episodes during deployment and therefore to better generalize to OOD situations. In the context of social navigation, we show that this improves the agent's capability to retain sparse, high-cost failures, such as human collisions. If the policy is trained in simulation, this also naturally addresses the sim-to-real gap, partially, by basing some of the decision making on real data. We integrate the explicit memory into a recurrent PPO architecture and use hidden states for memory retrieval to capture continuous spatiotemporal dynamics. The goal of exploiting rare, high-impact events is achieved by leveraging the RL agent's advantage signals. We train our agent in simulation with a combination of photorealistic rendering and non-visual crowd simulation and show that the agent is robust with respect to OOD social behavior.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑