用于密集不确定人群中机器人导航的人-人及人-机器人交互Transformer(H2INT)
Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds
浏览论文内容
中文总结 AI 辅助
本文提出H2INT强化学习框架,通过分层关系编码与课程学习提升密集不确定人群中机器人导航的安全性与鲁棒性,可迁移至不同人群布局,经仿真与真实机器人部署验证有效。
中文摘要 AI 辅助
密集人群中的机器人安全导航需要推理行人运动及其对机器人的响应变化。然而,许多基于学习的方法独立生成行人运动或假设均匀互惠,忽略了交互不确定性的重要来源。本文提出人-人及人-机器人交互Transformer(H2INT),这是一种强化学习框架,在策略学习期间保留行人运动中以机器人为条件的变化,同时允许不同行人的响应性存在差异。响应性在机器人可见时会影响人群动态,但不作为策略输入,策略必须从以机器人为中心的相对位置推断其后果。两阶段门控Transformer逐步编码人-人及人-机器人关系,循环策略捕捉其时间演化。课程学习逐步降低行人响应性以增加交互难度。仿真实验表明,在不同响应条件和人群密度下,与代表性基线相比,该方法的导航安全性和鲁棒性均有所提升,且无需重新训练即可迁移至结构不同的人群流动布局。消融实验验证了分层关系编码和门控更新的有效性,真实机器人部署进一步证实,所学策略可在物理环境中以稀疏观测运行。
英文摘要
Safe robot navigation in dense crowds requires reasoning about pedestrian motion and how it may change in response to a robot. However, many learning-based approaches generate pedestrian motion independently of the robot or assume uniform reciprocity, omitting an important source of interaction uncertainty. This paper presents a Human-Human & Human-Robot Interaction Transformer (H2INT), a reinforcement learning framework that retains robot-conditioned changes in pedestrian motion during policy learning while allowing responsiveness to vary across pedestrians. Responsiveness affects the crowd dynamics when the robot is visible but is not supplied as a policy input; the policy must instead infer its consequences from robot-centered relative positions. A two-stage gated Transformer progressively encodes human-human and human-robot relations, while a recurrent policy captures their temporal evolution. A curriculum gradually reduces pedestrian responsiveness to increase interaction difficulty. Simulation experiments demonstrate improved navigation safety and robustness over representative baselines across response conditions and crowd densities, and show transfer without retraining to structurally distinct crowd-flow layouts. Ablations support the hierarchical relational encoding and gated updates. Real-robot deployment further verifies that the learned policy can operate with sparse observations in a physical environment.
发表机构
- School of Automation, Beijing Institute of Technology(北京理工大学自动化学院)
- State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。