arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BodyCam-VQA:通过多模态推理与探针问题生成增强执法记录仪视频描述

BodyCam-VQA: Enhanced Body-Worn Camera Video Captioning via Multimodal Reasoning and Probe Question Generation

Karish Gupta, Matthew Alex, Alex Li, Yang Wu, Yun-Wei Chu, Kashif Munir, Xiaotian Zhou, Zhengping Ji, Xiaozhong Liu

arXiv 2609.10815首次发表:更新:

发表机构

Worcester Polytechnic Institute; Axon(伍斯特理工学院; Axon公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出自适应视觉问答框架BodyCam-VQA,通过结构化推理与探针问题生成,提升执法记录仪视频描述的细粒度取证细节提取能力,实现更可靠客观的执法记录。

AI 中文摘要

警方执法记录仪(BWC)录像已成为执法工作的关键组成部分,它确保了法律透明度、警察问责制以及公民权利的保护。然而,由于此类数据的多模态视频格式,有效处理这些数据仍然是一项重大挑战。在许多情况下,BWC视频包含混乱的场景、低视觉质量、快速移动/互动以及高噪声音频,这使得即使是目前最先进的多模态模型也难以进行视觉理解。当前的视觉语言模型(VLM)经常忽略关键的取证细节,例如有价值证据的存在或嫌疑人与警察互动的潜在细微差别,而这些对于公正的法律结果以及平民/警察安全至关重要。为解决这些局限性,我们提出了一种专为高风险执法场景设计的自适应视觉问答(VQA)框架。该框架采用结构化推理方法来提取传统描述系统无法捕获的细粒度视觉证据。我们尝试了多种问题生成模型,包括基础模型和微调的开源权重模型,以观察不同问题生成模型实现之间的性能差异。我们的结果表明,这种基于VQA的架构能够提供更可靠、客观且详细的执法事件记录,最终通过人工智能辅助的取证清晰度,成为保护执法人员和公众的有力工具。

英文摘要

Police body-worn camera (BWC) footage has emerged as a critical aspect of law enforcement that ensures legal transparency, officer accountability, and the protection of civil rights. However, effectively processing this data remains a significant challenge due to its multimodal video format. BWC videos, in many cases, comprise chaotic scenes with low visual quality, rapid movement/interactions, and high-noise audio that make visual understanding a challenge for even SOTA multimodal models. Current Vision-Language Models (VLMs) frequently overlook critical forensic details, such as the presence of valuable evidence or the latent nuances of suspect-officer interactions, which are vital for fair legal outcomes and civilian/officer safety. To address these limitations, we propose an Adaptive Visual Question Answering (VQA) framework engineered for high-stakes law enforcement. Our framework employs a structured reasoning approach to extract fine-grained visual evidence that traditional captioning systems fail to capture. We experiment with multiple question generation models, including foundation models and fine-tuned open-weight models, to observe performance variation among question generation model implementations. Our results demonstrate that this VQA-driven architecture provides a more reliable, objective, and detailed record of enforcement events, ultimately serving as a powerful tool to protect both law enforcement officers and the public through AI-assisted forensic clarity.

CommentsEMNLP 2026 Workshop NLP4PI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑