表示-行为对齐用于可解释的弱监督视频异常检测
Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection
- Sun Yat-sen University(中山大学)
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多模态大语言模型在弱监督视频异常检测中表示与行为错位的问题,提出参数高效的表示-行为对齐方法,仅更新0.012%参数,提升决策方向与判别表示的匹配,实现可解释检测。
AI中文摘要:
多模态大语言模型(MLLMs)为视频异常检测提供了使其更具可解释性的自然途径。然而,它们的最终决策并不总是充分利用其隐藏状态中包含的判别性信息,我们将这一问题称为表示-行为错位。我们将这一差距分解为两部分:容量组件,用于衡量从未聚合到读出位置的判别性信息;以及方向组件,用于衡量该位置最优轴与原生正常-异常轴之间的角度失配。在多个视频异常检测基准和MLLM骨干网络上,方向组件占主导地位,残差流追踪显示,在多个中后期注意力层中,原生轴的可分性急剧上升。由于这两个组件均由注意力而非MLP更新所控制,我们提出了表示-行为对齐(RBA),这是一种参数高效的方法,仅使用视频级标签来调整这些层,同时更新约0.012%的骨干网络参数。在三个基准上的实验表明,RBA提升了原生读出性能,并更好地将模型的决策方向与判别性表示对齐,且通过单一生成过程产生异常决策和解释。
英文摘要:
Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012\% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.