AI 中文总结
研究开放世界视频异常检测中定义盲目性问题,分解动态定义评估,引入三个定义条件评估指标及DeCoS,指出OWVAD应按定义条件异常评分评估,改进了多个基线模型的定义跟随能力。
AI 中文摘要
开放世界视频异常检测(OWVAD)旨在检测符合用户指定异常定义的事件,这一要求比一般异常定位更强。当前OWVAD评估很大程度上未能分离这种条件行为,标准VAD指标和动态定义协议可能被目标与正常分离主导,导致模型对查询定义不敏感,即定义盲目性。通过分解动态定义评估,发现目标与正常检测权重更大。为此引入三个定义条件评估指标,实验表明多个基线定位异常时刻但定义跟随能力弱。还引入DeCoS,改进了最强基线。结果表明OWVAD应按定义条件异常评分评估,而非不同提示标签下的异常检测。
英文摘要
Open-world video anomaly detection (OWVAD) aims to localize anomalous events in videos according to a user-specified definition. An OWVAD detector should therefore use the supplied definition to decide what is anomalous. We call this behavior definition following, and its failure definition blindness. However, existing temporal evaluation protocols may obscure this distinction, making strong generic anomaly detection appear to reflect definition following. For example, Drift@5 and the customizable VAD protocol treat both normal frames and anomalies outside the current definition as negatives. Across the three benchmarks used in our study, normal frames account for 78.6%-96.4% of these negatives. As a result, the reported score can be driven largely by target-versus-normal separation rather than target-versus-other-anomaly separation. Consistent with this diagnosis, a LaGoVAD branch that never receives the anomaly definition nearly matches the full model under Drift@5. We therefore introduce three complementary definition-conditioned evaluations that separate definition following from generic anomaly detection. Applying them to VAD systems and general vision-language models on UCF-Crime, XD-Violence, and MSAD reveals that strong performance under existing evaluation can coexist with weak definition following. Motivated by these results, we introduce Definition-Contrastive Scoring (DeCoS), which reduces anomaly evidence shared across competing definitions and improves definition following across the three benchmarks. Our results show that generic anomaly detection and definition following are distinct capabilities and should be evaluated separately in OWVAD.
CommentsPreprint