发表机构
Lehigh University(理海大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Social-WM提出一种潜在世界模型规划框架,通过动作条件化预测和逆动力学目标学习社交导航中的安全约束,在Social-HM3D上实现63.77%成功率并显著降低碰撞率,且支持零样本迁移。
AI 中文摘要
安全的社交导航要求机器人不仅能够预测其行为未来的后果,还要预测在周围物理和社交约束下,名义上的行为是否真的能够被执行。我们提出了Social-WM,一个从以自我为中心的RGB视频序列中训练的高效潜在世界模型规划框架。我们的关键观察是,社交导航经验中存在着名义行为与可实现行为之间的系统性差异:名义上的前进行为在自由空间中可能完全执行,但当朝向行人或障碍物时则需要受到约束。Social-WM直接通过动作条件化的未来预测来学习这些与安全相关的后果,其中目标是每条命令之后实际观察到的未来状态。我们进一步引入了一个可实现的逆动力学目标,该目标将观察到的潜在状态转移与实际实现的动作而非名义动作相关联。在部署时,候选动作通过潜在世界模型进行想象,逆动力学模型估计其可实现性;名义-可实现差异在执行前提供了安全信号。学习到的动力学和可实现性模型保持目标无关,并支持基于位置和图像目标的导航。在Social-HM3D上,Social-WM实现了63.77%的成功率,同时将人类碰撞率降低到21.67%,并在零样本迁移到Social-MP3D时保持强劲性能,无需显式行人跟踪、特权人类状态或在线强化学习。
英文摘要
Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.
Comments9 pages, 5 figures