发表机构
National Chung Hsing University; Toyota Technological Institute(国立中兴大学; 丰田工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PRISM通过被动观察推断行人交互特质,利用序数潜在空间和时间稳定性评分,在人群导航中降低碰撞率并小幅提升效率。
AI 中文摘要
人类在互动前通常会观察他人,并据此调整自身行为。然而,机器人在人群中的导航往往主要依据观测到的几何状态来表征行人,从而隐含地忽略了互动倾向中的个体差异。我们提出了PRISM(交互风格与运动的预测性表征),这是一个从人类互动的被动观察中推断交互特质的框架。PRISM利用由Rank-N-Contrast损失训练的Transformer编码器,将人类轨迹编码到一个连续的序数潜在空间中,并将每个推断出的特质与提供给导航策略的时间稳定性评分配对。在随机人群模拟中,与仅基于几何信息的基线相比,PRISM降低了碰撞率,并在导航时间和路径长度指标上带来了小幅改进。这些结果表明,在动态人群中,被动潜在特质推断对社交导航具有实用性。
英文摘要
Humans often observe others before interacting and adjust their behavior accordingly. Robot navigation in crowds, however, often represents pedestrians mainly by observed geometric states, leaving individual differences in interaction tendencies implicit. We propose PRISM (Predictive Representation of Interaction Style and Motion), a framework that infers interaction traits from passive observations of human-human interactions. PRISM encodes human trajectories into a continuous ordinal latent space with a transformer encoder trained by Rank-N-Contrast loss, and pairs each inferred trait with a temporal-stability score supplied to the navigation policy. In randomized crowd simulations, PRISM reduces collision rates over the geometry-only baseline and yields small improvements in navigation-time and path-length metrics. These results suggest the utility of passive latent-trait inference for social navigation in dynamic crowds.
CommentsECCV 2026 Workshop on Agent in World