STARS:从时空动态到人机交互中的社会表征
STARS: From Spatiotemporal Dynamics to Social Representations in Human-Robot Interaction
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出SocialNav-SUB基准,评估视觉语言模型在社交机器人导航中的场景理解能力,发现其仍逊于简单规则和人类共识,揭示了关键差距。
AI中文摘要:
在动态、以人为中心的环境中进行机器人导航,需要基于稳健的场景理解做出符合社会规范的行为决策。最近的视觉语言模型(VLMs)展现出有前景的能力,如物体识别、常识推理和上下文理解,这些能力与社会机器人导航的细致需求相契合。然而,目前尚不清楚VLMs能否准确理解复杂的社会导航场景(例如,推断智能体之间的时空关系及人类意图),而这对于安全且符合社会规范的机器人导航至关重要。尽管一些近期工作探索了VLMs在社会机器人导航中的应用,但尚无现有工作系统性地评估它们满足这些必要条件的能力。在本文中,我们引入了社会导航场景理解基准(SocialNav-SUB),这是一个视觉问答(VQA)数据集和基准,旨在评估VLMs在真实世界社会机器人导航场景中的场景理解能力。SocialNav-SUB提供了一个统一框架,用于在需要空间、时空和社会推理的VQA任务中,将VLMs与人类和基于规则的基线进行比较。通过对最先进的VLMs进行实验,我们发现,尽管性能最佳的VLM在同意人类答案方面达到了令人鼓舞的概率,但它仍不如更简单的基于规则的方法和人类共识基线,这表明当前VLMs在社会场景理解方面存在关键差距。我们的基准为社会机器人导航的基础模型进一步研究奠定了基础,提供了一个框架来探索如何定制VLMs以满足真实世界的社会机器人导航需求。本文的概述以及代码和数据可在以下https URL找到。
英文摘要:
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear whether VLMs can accurately understand complex social navigation scenes (e.g., inferring the spatial-temporal relations among agents and human intentions), which is essential for safe and socially compliant robot navigation. While some recent works have explored the use of VLMs in social robot navigation, no existing work systematically evaluates their ability to meet these necessary conditions. In this paper, we introduce the Social Navigation Scene Understanding Benchmark (SocialNav-SUB), a Visual Question Answering (VQA) dataset and benchmark designed to evaluate VLMs for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across VQA tasks requiring spatial, spatiotemporal, and social reasoning in social robot navigation. Through experiments with state-of-the-art VLMs, we find that while the best-performing VLM achieves an encouraging probability of agreeing with human answers, it still underperforms simpler rule-based approach and human consensus baselines, indicating critical gaps in social scene understanding of current VLMs. Our benchmark sets the stage for further research on foundation models for social robot navigation, offering a framework to explore how VLMs can be tailored to meet real-world social robot navigation needs. An overview of this paper along with the code and data can be found at https://larg.github.io/stars.