发表机构
Interactive Robotics and Language Lab, University of Maryland Baltimore County; DEVCOM Army Research Laboratory(马里兰大学巴尔的摩县分校交互式机器人与语言实验室; 陆军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在高风险环境中,提出基于VR的人机交互框架,利用VLM辅助机器人识别潜在危险并标注兴趣点,通过VR界面呈现给操作员,经实验验证该方法能有效支持态势感知,提升交互体验。
AI 中文摘要
在诸如灾难响应等高风险环境中,态势感知不仅依赖于检测危险,还取决于将危险清晰传达给人类操作员。视觉语言模型(VLM)在安全关键场景的场景理解中显示出强大潜力,但其作为面向人类的机器人系统一部分的价值仍未得到充分探索。我们提出了一个基于虚拟现实的人机交互框架,用于研究VLM辅助机器人如何在模拟危险环境中支持态势感知。在我们的系统中,机器人探索虚拟场景并查询VLM以识别潜在危险并标注面向用户的兴趣点。这些标注通过沉浸式虚拟现实界面呈现给人类操作员。该框架能够对机器人危险识别以及向用户传达安全关键信息进行可控评估。我们的研究结果表明,带标注的虚拟现实界面优于无标注的基线,并且参与者在与系统交互时报告了高清晰度、有用性和舒适度。这些发现表明,将基于VLM的机器人感知与沉浸式可视化相结合是在危险环境中支持态势感知的一种有前途的方法。
英文摘要
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating them clearly to human operators. Vision Language Models (VLMs) have shown strong potential for scene understanding in safety-critical settings, yet their value as part of human-facing robotic systems remains underexplored. We present a VR-based Human Robot Interaction framework for studying how VLM-assisted robots can support situational awareness in simulated hazardous environments. In our system, a robot explores a virtual scene and queries a VLM to identify potential hazards and annotate user-facing points of interest. These annotations are presented to a human operator through an immersive VR interface. This framework enables controlled evaluation of both robotic hazard identification and the communication of safety-critical information to users. Results from our study indicate that the annotated VR interface was preferred over the unannotated baseline and that participants reported high clarity, usefulness, and comfort when interacting with the system. These findings suggest that combining VLM-based robotic perception with immersive visualization is a promising approach for supporting situational awareness in hazardous settings.
Comments7 Pages, Accepted to RO-MAN 2026