arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

定位去偏:弱监督空间异常检测的补丁级基准与基线

Localizing to Debias: A Patch-Level Benchmark and Baseline for Weakly Supervised Spatial Anomaly Detection

Sara Abdulaziz, Abdulrahman Al-Abri, Giacomo D'Amicantonio, Egor Bondarev

arXiv 2608.12045首次发表:更新:

发表机构

Eindhoven University of Technology(埃因霍温理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对弱监督视频异常检测中存在的背景偏差问题,提出SST-WSVADL稀疏时空框架,结合动态稀疏化与运动感知正则化,发布相关数据集和评估协议,实现异常定位与场景偏差审计,性能与现有方法相当。

AI 中文摘要

尽管弱监督视频异常检测(WSVAD)的研究日益增多,但现有方法仍难以弥合粗粒度时间监督与细粒度空间推理之间的差距。一个关键障碍是,时间检测器倾向于依赖背景和场景级线索,而非真正具有判别性的异常证据。这种背景偏差引发了伦理问题:模型可能会无意中将异常与社会或环境背景关联,而非与真实的犯罪相关线索关联。若缺乏空间定位,此类偏差将保持隐蔽且无法审计。为解决这一问题,我们提出SST-WSVADL,这是一种将时间异常检测与细粒度空间定位相结合的稀疏时空框架。SST-WSVADL不会不加区分地处理所有空间区域,而是通过动态稀疏化逐步聚焦于最具异常相关性的时空区域,自然抑制背景主导内容,同时保留判别性证据。时间分支与空间分支通过运动感知正则化端到端耦合,该正则化引导稀疏化朝向动态信息丰富的区域,无需依赖外部检测器或视觉语言提示。我们公开发布了针对三个公共数据集(UCF-Crime、XD-Violence和MSAD)的帧级空间注释以及与方法无关的评估协议。这些资源使研究社区能够审计WSVAD预测中的空间偏差,推动实现更具伦理性和可问责性的异常检测。实验表明,SST-WSVADL在各基准上与现有方法具有竞争力,同时实现了场景偏差的定位和补丁级可审计性,为面向可解释性的WSVAD模型评估提供了可复现的基础。

英文摘要

Despite growing interest in weakly supervised video anomaly detection (WSVAD), current methods struggle to bridge the gap between coarse temporal supervision and fine-grained spatial reasoning. A key obstacle is the tendency of temporal detectors to latch onto background and scene-level cues rather than truly discriminative anomaly evidence. This background bias raises ethical concerns: models may inadvertently associate anomalies with societal or environmental context rather than authentic crime-related cues. Without spatial grounding, such biases remain hidden and unauditable. To address this, we propose SST-WSVADL, a sparse spatio-temporal framework that bridges temporal anomaly detection with fine-grained spatial localization. Rather than processing all spatial regions indiscriminately, SST-WSVADL progressively focuses on the most anomaly-relevant spatio-temporal regions through dynamic sparsification, naturally suppressing background dominant content while preserving discriminative evidence. The temporal and spatial branches are coupled end-to-end via motion-aware regularization that guides sparsification toward dynamically informative regions, without relying on external detectors or vision-language prompts. We publicly release frame-level spatial annotations and a method-agnostic evaluation protocol for three public datasets: UCF-Crime, XD-Violence, and MSAD. These resources enable the community to audit spatial biases in WSVAD predictions, supporting progress toward more ethical and accountable anomaly detection. Experiments demonstrate that SST-WSVADL is competitive with prior methods across benchmarks while enabling localization and patch-level auditability of scene bias, providing a reproducible foundation for interpretability-oriented evaluation of WSVAD models.

CommentsECCV 2026 FAILED Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑