MuST-VAD:用于视频异常检测的互结构化学习
MuST-VAD: Mutual Structured Learning for Video Anomaly Detection
中文总结 AI 辅助
MuST-VAD是用于弱监督视频异常检测的互结构化学习框架,通过检测器与LVLM的双向知识交换,在UCF-Crime数据集上显著提升了检测精度,AP优于现有最优方法。
中文摘要 AI 辅助
本文提出了MuST-VAD,一种用于弱监督视频异常检测(VAD)的互结构化学习框架,其中异常检测器和大型视觉语言模型(LVLM)交换各自获取的知识。弱监督VAD中的检测器从固定的、与任务无关的主干提取的特征中学习异常分数,这些固定特征限制了可达到的检测精度。因此,近期方法将LVLM语义作为更丰富的特征迁移到检测器中,但这种迁移是单向的:检测器从目标视频中学到的内容不会反馈给LVLM。MuST-VAD将单向迁移扩展为双向学习循环,在此循环中,最新的检测器预测监督LVLM的适配,适配后的LVLM返回更新后的表示以重新训练检测器;两个模型在小型视频组上交替进行这些更新。两个模型均在检测器选定的关键片段上训练,同时置信度加权和基于注释的问答确保交换的监督可靠。在UCF-Crime数据集上,我们的互学习将单向迁移基线的AUROC从88.15%提升至88.63%,平均精度(AP)从37.25%提升至42.46%,在AP指标上优于现有最优方法4.13个百分点。
英文摘要
In this paper, we propose MuST-VAD, a mutual structured learning framework for weakly supervised video anomaly detection (VAD) in which an anomaly detector and a large vision-language model (LVLM) exchange their acquired knowledge. Detectors in weakly supervised VAD learn anomaly scores from features extracted by a fixed, task-agnostic backbone. These fixed features bound the achievable detection accuracy. Recent methods therefore transfer LVLM semantics into the detector as richer features. However, this transfer is one-way: what the detector learns about the target videos never returns to the LVLM. MuST-VAD extends the one-way transfer into a bidirectional learning loop. In this loop, the latest detector predictions supervise the LVLM adaptation, and the adapted LVLM returns updated representations that retrain the detector; the two models alternate these updates over small video groups. Both models train on detector-selected key clips, while confidence weighting and annotation-anchored question answering keep the exchanged supervision reliable. On UCF-Crime, our mutual learning improves the one-pass transfer baseline from 88.15% to 88.63% AUROC and from 37.25% to 42.46% average precision (AP), outperforming the state-of-the-art method in AP by 4.13 points.