arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27327cs.CVcs.HC

视觉语言模型能否分析以人为中心的视频?映射模型能力与人机协作工作流

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

  • University of Washington(华盛顿大学)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang, Huanfen Yao, Shwetak Patel, Zhihan Zhang, Jacob O. Wobbrock

AI总结:

本研究通过构建五维分类体系和15个任务基准,系统评估视觉语言模型在分析以人为中心视频中的能力,发现模型单独标注接近人类水平,而人类验证工作流在提升准确率的同时大幅降低时间和成本。

AI中文摘要:

视频为人类行为、互动和情境背景提供了丰富的记录,为理解人类和开展以人为中心的研究提供了重要证据。随着视觉语言模型(VLMs)在视频分析方面能力日益增强,它们为自动化这一传统上高度依赖人工的过程提供了机会。然而,一个核心问题依然存在:视觉语言模型何时能够独立分析以人为中心的视频,何时可靠的分析仍然需要人类参与?为解决这一问题,我们首先刻画了以人为中心研究中视频分析实践的现状。我们系统分析了所有1,702篇CHI 2026全文论文,识别出125篇对视频进行标注的论文。通过迭代编码,我们构建了一个涵盖分析目的、视角、现象、推理需求和标注权威性五个维度的分类体系。基于该分类体系所捕获的重复出现的标注任务,我们从开放数据集中构建了一个包含15个代表性任务的基准,以映射通用视觉语言模型的能力与局限。通过比较三种标注工作流:仅视觉语言模型、仅人类、以及人类对视觉语言模型输出的验证,我们考察了人类与视觉语言模型之间的分工。在各项任务中,仅视觉语言模型的标注在平均准确率上接近人类水平(HNS = 97.0,其中100表示仅人类表现的得分),展示了自动化以人为中心视频分析的巨大潜力。人类验证实现了最高准确率(HNS = 121.5),同时相对于仅人类标注,将人类标注时间减少了48.9%,货币成本降低了31.3%-44.5%。我们的研究结果将现实世界中以人为中心的视频分析任务与当前视觉语言模型的能力联系起来,并阐明了人机协作如何使视觉语言模型辅助的分析既可靠又高效。

英文摘要:

Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.

↑