arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16081cs.CV

SafeGesture:通过场景条件安全解释评估视觉语言模型的细粒度手势理解能力

SafeGesture: Evaluating Fine-Grained Hand Gesture Understanding in Vision-Language Models through Scenario-Conditioned Safety Interpretation

Taegang Kim, Saleh Afroogh, Junfeng Jiao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SafeGesture基准,评估5款视觉语言模型的细粒度手势安全理解能力,发现模型存在感知与推理脱节,瓶颈为场景条件安全推理而非手势识别。

中文摘要 AI 辅助

开源权重的前沿视觉语言模型(VLMs)在通用图像理解方面表现良好,但其在安全关键操作场景中解释细粒度手势的能力仍未得到充分研究。本文提出SafeGesture基准,用于评估模型是否能从手势中推断出符合场景的安全动作,该基准将6种HaGRID手势与8种操作场景配对,共包含4800个样本,并对Qwen2.5-VL-7B、LLaVA-NeXT-7B、InternVL2-8B、Phi-3.5-Vision和GPT-4o这5个模型进行评估。结果揭示了感知与推理的脱节:GPT-4o的手势准确率达98.4%,但安全准确率仅为53.3%;Qwen2.5-VL的对应数值为84.9%和39.5%,二者的差距分别为45.0和45.4个百分点。5个模型中有4个很少或从未使用不确定性标签,且各模型的失败方向差异显著。准确率还掩盖了标签偏差:一项无视觉输入的场景多数策略达到了58.3%,高于所有被评估的模型,而仅GPT-4o在宏F1指标下超过该基线。视觉输入使安全准确率提升了11.2至30.2个百分点,但提供真实手势文本仅使性能提升0.4至3.2个百分点,且无模型超过56.2%。这些结果表明,主要瓶颈在于场景条件安全推理而非手势识别。

英文摘要

Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained hand gestures in safety-critical operational contexts remains largely unexamined. We introduce SafeGesture, a benchmark that evaluates whether a model can infer scenario-appropriate safety actions from hand gestures. It pairs six HaGRID gestures with eight operational scenarios for 4,800 items and evaluates Qwen2.5-VL-7B, LLaVA-NeXT-7B, InternVL2-8B, Phi-3.5-Vision, and GPT-4o. Results reveal a perception-reasoning decoupling: GPT-4o achieves 98.4% gesture accuracy but 53.3% safety accuracy, while Qwen2.5-VL reaches 84.9% and 39.5%, yielding gaps of 45.0 and 45.4 percentage points. Four of five models rarely or never use the uncertainty label, and failure directions differ substantially across models. Accuracy also obscures label bias: a scenario-majority policy with no visual input reaches 58.3%, above every evaluated model, while only GPT-4o exceeds this prior under macro-F1. Visual input improves safety accuracy by 11.2 to 30.2 percentage points, but providing the ground-truth gesture as text improves performance by only 0.4 to 3.2 points, and no model exceeds 56.2%. These results indicate that the main bottleneck is scenario-conditioned safety reasoning rather than gesture recognition.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Urban Information Lab, The University of Texas at Austin(德克萨斯大学奥斯汀分校城市信息实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑