arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37476cs.ROcs.CVcs.LG

在策略状态空间中从互联网视频学习社交导航

Learning Social Navigation from Internet Videos in the Policy State Space

  • National University of Singapore(新加坡国立大学)
  • Singapore Management University(新加坡管理大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • Max Planck Institute for Plasma Physics(马克斯·普朗克等离子体物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna, Diwen Liu, Zhengcheng Shen, Harold Soh

AI总结:

提出从互联网单目视频构建社交导航训练环境的流程,在策略状态空间直接模拟,使策略在Arena基准上达81.2%成功率并成功完成19/20真实试验。

AI中文摘要:

训练稳健的社交导航策略需要具备多样化场景布局、地形和人类运动的模拟器,但构建此类环境并指定行人行为成本高昂。我们提出了一种高效流程,将普通的单目步行视频直接转换为策略状态空间中的闭环社交导航训练环境。我们的关键观察是,局部社交导航主要依赖两类信息:机器人可通行区域以及附近行人的移动方式。因此,我们将静态场景表示为度量可通行性地图,该地图可在反事实机器人运动下进行刚性变换,同时随时间直接回放从视频中恢复的行人轨迹。这种抽象使我们能够直接在策略状态空间中定义前向动力学,并高效模拟反事实机器人状态,而无需重建或渲染逼真观测。所得策略在独立Arena基准上达到81.2%的成功率,而最强基线为75.0%,并且在无需策略微调的情况下,在20次真实机器人试验中成功19次。

英文摘要:

Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning. Project page: https://jiaming.im/VideoSocNav

补充信息

↑