无解析器VLM验证用于联邦弱监督视频异常检测
Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection
浏览论文内容
中文总结 AI 辅助
针对联邦弱监督视频异常检测,提出轻量级MIL-VLM级联框架,利用冻结VLM的logit接口验证可疑片段,无需解析器,在UCF-Crime上提升帧级AUC和AP。
中文摘要 AI 辅助
视觉语言模型如何在监控数据保持分布式、弱标注且资源受限的情况下帮助视频异常检测(VAD)?大多数弱监督VAD方法假设集中式训练;近期基于VLM的扩展进一步依赖密集推理、生成的解释或额外适配。我们提出一种轻量级联邦MIL-VLM级联框架,其中仅跨客户端训练一个紧凑的MIL评分器,而冻结的VLM事后验证高评分可疑片段。我们研究了两种VLM反馈接口:解析的文本生成决策和基于logit的接口,后者从下一个词元是/否概率中提取连续异常分数。在UCF-Crime数据集上使用InternVL3.5-2B和Qwen3-VL-2B-Instruct进行的实验表明,文本生成验证在诊断性时间后处理后能提高帧级AUC,但仍对提示、解析器、模型选择和平滑敏感。相比之下,logit接口提供固定的无解析器信号,在两种VLM上均比MIL基线提高了帧级AUC和帧级AP,且其主要配置无需时间后处理。由于可疑片段一旦可用即独立更新,下一个词元logit反馈为文本生成验证提供了一种简单的片段局部替代方案。
英文摘要
How can vision-language models help video anomaly detection (VAD) when surveillance data remain distributed, weakly labeled, and resource-constrained? Most weakly supervised VAD methods assume centralized training; recent VLM-based extensions further rely on dense inference, generated explanations, or additional adaptation. We introduce a lightweight federated MIL-VLM cascade in which only a compact MIL scorer is trained across clients, while a frozen VLM verifies high-scoring suspect segments post hoc. We study two VLM feedback interfaces: parsed text-generation decisions and a logit-based interface that extracts a continuous anomaly score from next-token Yes/No probabilities. Experiments on UCF-Crime with InternVL3.5-2B and Qwen3-VL-2B-Instruct show that text-generation verification can improve frame-level AUC after diagnostic temporal post-processing, but remains sensitive to prompts, parsers, model choice, and smoothing. In contrast, the logit interface provides a fixed parser-free signal that improves both frame-level AUC and frame-level AP over the MIL baseline across both VLMs, without temporal post-processing in its main configuration. Since suspect segments are updated independently once available, next-token logit feedback provides a simple segment-local alternative to text-generation verification.
发表机构
- ESIEA
- CY Cergy Paris University(CY塞尔吉-巴黎大学)
机构由 AI 辅助整理,请以论文原文为准。