arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21854cs.CV

弱监督视频异常检测中的帧级评估大多衡量的是视频级排名

Look Inside Each Video: Rethinking How Video Anomaly Detection Is Evaluated

Inpyo Song, Jangwon Lee

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过分解Micro-AUROC指标发现,弱监督视频异常检测的帧级评估实际多衡量视频级排名,检测器可通过视频级分类实现高帧级评分,无需准确排序视频内时刻。

中文摘要 AI 辅助

弱监督视频异常检测器使用视频级标签进行训练,但通常作为时间定位器使用Micro-AUROC或对汇集的测试帧计算AP进行评估。由于这些指标比较来自不同视频的帧,检测器可以通过区分视频而无需准确对视频内的时刻进行排序来获得高分。我们按视频身份对Micro-AUROC进行精确分解,得到用于视频内时间排序的Within-AUROC和用于视频间比较的Cross-AUROC。在ShanghaiTech、XD-Violence和UCF-Crime数据集上,异常帧与正常帧之间的比较中仅有0.071%-0.388%发生在同一视频内。当两类样本仍分布在V个视频中时,该比例以O(1/V)的速率下降,我们将这一基准特性称为时间稀释。我们在相同的视频级监督下训练异常视频二分类器,并将每个视频的分数重复到所有帧上。这些无视频内变化的视频恒定输出仍能达到81.40%-97.18%的Micro-AUROC。在72次受控运行中,将每个帧分数替换为其视频均值可保留高于随机水平的Micro-AUROC余量的中位数98.6%。相同的经验模式也适用于作者发布的输出,以及XD-Violence在其官方AP评估下的情况。因此,检测器即使对每个视频内的所有时刻都赋予相同的分数,也能达到较高的汇集分数。

英文摘要

Video anomaly detection (VAD) aims to localize anomalous events by identifying when they occur. A common evaluation pools frame-level anomaly scores across all test videos and measures whether anomalous frames receive higher scores than normal frames (Micro-AUROC). This pooling creates comparisons both within the same video and across different videos. However, the pooled evaluation protocol commonly used in current VAD benchmarks makes Micro-AUROC heavily dependent on cross-video comparisons, which do not directly test whether anomalous frames rank above normal frames within the same video. To quantify this imbalance, we decompose Micro-AUROC into within-video and cross-video comparisons. Across four widely used weakly supervised VAD benchmarks, within-video pairs account for only 0.071-0.388% of all anomalous-normal frame pairs. This imbalance can make Micro-AUROC a poor indicator of within-video ranking. Video-level classifiers that assign a single score to each video achieve 81.40-97.18% Micro-AUROC while obtaining 50% Within-AUROC. For existing weakly supervised VAD detectors, replacing all frame scores in each video with their mean sets Within-AUROC to 50%. Across 120 runs, this replacement retains a median 98.8% of the original Micro-AUROC margin above chance. Because the decomposition depends on the evaluation protocol rather than the training supervision, the same distinction applies whenever VAD frame scores are pooled across videos. High Micro-AUROC alone therefore does not establish temporal localization. Our analysis reveals a limitation of pooled VAD evaluation and motivates reporting within-video ranking alongside Micro-AUROC.

发表机构

  • SungKyunKwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑