视觉-语言模型能否从机器人的第一人称视角图像评估人际距离风险?
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
- National University of Kyiv-Mohyla Academy(基辅莫希拉国立大学)
- University of Turin(都灵大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究评估了InternVL、Qwen-VL、SmolVLM三种VLM在机器人第一人称视角图像人际距离危险分类中的表现,发现Qwen-VL经特定优化后高危险召回率更高,当前VLM在相关推理定位上仍有局限。
AI中文摘要:
从机器人的第一人称视角评估人际距离危险,对于在人类环境中实现安全的具身导航至关重要,该任务需要视觉与上下文推理能力。我们评估了三个开源视觉-语言模型(VLM):InternVL、Qwen-VL和SmolVLM,将机器人第一人称视角图像分类为四个危险等级,对比了三种提示策略和两轮QLoRA微调,与分层随机基线进行比较。未进行微调时,所有模型的表现接近基线;微调仅带来适度的整体改进,但采用高级提示的Qwen-VL对高危险案例的召回率显著高于其他模型。对人物定位的进一步分析显示,正确的危险分类并不对应更好的空间定位,这表明模型可能无需关注场景的相关区域就能生成有用的安全标签。这些结果表明,当前VLM在细粒度人际距离推理和空间定位方面仍存在局限,不过针对性的提示和微调可提升选定模型的高危险检测能力。
英文摘要:
Assessing proxemic danger from a robot's egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (\textit{InternVL}, \textit{Qwen-VL}, and \textit{SmolVLM}) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tuning against a stratified random baseline. Without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. However, \textit{Qwen-VL} with an advanced prompt achieves substantially higher recall for high-danger cases than the other models. An analysis of person localization further shows that correct danger classification does not correspond to better spatial grounding, indicating that a model may produce a useful safety label without attending to the relevant region of the scene. These results show that current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, although targeted prompting and fine-tuning can improve high-danger detection in selected models.