VLM 的回答并非异常分数:无训练视频异常检测中的排名压缩
A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection
查看机构详情
- SungKyunKwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对基于VLM的无训练视频异常检测,发现生成式回答读出存在排名压缩问题,提出概率读出方法可提升性能,证明回答接口是该类检测器的关键组成部分。
中文摘要 AI 辅助
视觉语言模型(VLM)可通过回答视频片段相关问题实现无训练视频异常检测(VAD)。然而,VAD基准要求每个片段输出标量异常分数,并使用AUROC或AP评估所得排名。因此,基于VLM的检测器应定义一个回答接口:回答尺度指定可接受的回答,读出规则将模型的输出分布映射为分数。由于该接口会改变评估排名,它属于检测器的一部分,而非格式细节。生成式读出仅使用最可能的回答,而概率读出则利用了所有可接受回答的完整分布。在四个7-8B规模的VLM上,对于每一组测试的回答尺度、基准和指标组合,概率读出均优于生成式读出,在四个基准-指标对上的平均提升幅度为5至13个百分点。这种差距的产生是因为生成式读出仅保留每个片段的一个回答值,导致具有不同回答分布的片段可能获得相同分数并丢失相对顺序,我们将这种相对顺序的丢失称为生成式回答排名压缩。即使回答尺度允许91个回答,生成式读出也仅产生4至18个不同分数,而概率读出则保留了精细得多的分数分辨率。我们测试的每一种解码策略、提示措辞以及联合评分-解释提示下,该优势均持续存在。因此,回答接口是基于VLM的VAD中具有重要影响的组成部分,应明确指定并进行评估。
英文摘要
Anomaly detection aims to identify observations that deviate from normal patterns. Recent work has increasingly used pretrained vision--language models (VLMs) for training-free image and video anomaly detection without task-specific retraining. Anomaly detection is commonly evaluated by how well anomaly scores rank anomalous images or video frames above normal ones. Generative VLMs, however, assign probabilities to possible answers and then decode a single answer. This decoding step may discard ordering information. We call this loss decoded-answer rank compression. To isolate this effect, we compare two ways of scoring the same VLM output: one uses only the decoded answer, while the other computes a probability-weighted score over all possible answers. Probability-weighted scoring consistently outperforms decoded-answer scoring across image and video anomaly detection benchmarks, using different VLMs and answer scales. The mean gains range from 7.66 to 19.95 points on the primary metrics. Most of this gap comes from decoded-answer ties. Breaking decoded-answer ties with answer probabilities recovers at least 95% of the average performance gap on every benchmark. We further investigate how these ties affect evaluation. We find that evaluating tied scores one input at a time makes the reported AUROC depend on input order, and that linear PR interpolation can inflate reported performance. These results reveal that both how VLM outputs are scored and how discrete anomaly scores are evaluated affect reported anomaly-detection performance.