arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27879cs.CVcs.LG

交互表示究竟测量什么?弱监督暴力检测中的事件前可分性

What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection

  • Auburn University(奥本大学)

机构由 AI 辅助整理,请以论文原文为准。

Parishruthi Ganesh

AI总结:

该研究在弱监督暴力检测中对比不同交互表示的性能,发现事件前来源线索会掩盖表示差异,提出的诊断方法仅需基准自带标注,揭示了视频级AUC的构成。

AI中文摘要:

关节人体姿态提供了超出粗略空间关系的详细身体配置信息,但当下游管道固定时,这种细节是否能产生更具区分性的信息仍不清楚。我们通过早期暴力检测来研究这一问题。在固定跟踪器、时间头、监督、折叠和评估的情况下,我们比较了五种交互表示,涵盖粗略边界框几何、匹配的手工姿态类似物、丰富的姿态描述符以及从原始关节学习的容量匹配编码器,采用带聚类自举区间的视频级评估。在包含15个异常视频的该子集上,没有基于姿态的表示优于粗略几何,但无法排除存在小效应的可能。将管道扩展到冻结的视觉编码器,并在XD-Violence(137个异常视频,是我们UCF-Crime样本的9倍)上重复比较,人物裁剪外观和全帧上下文均大幅优于几何,不过在UCF-Crime上上下文与外观表现相当,在更大的划分上上下文优于外观:裁剪到交互人物并不比编码全帧更有优势。这促使我们直接测试基准测量的内容。在移除序列长度作为线索的控制条件下,仅使用标注 onset 之前的帧对异常视频评分,在两个基准上保留了39-91%的高于随机的可分性,包括7个手工设计的几何通道。对最紧凑的 onset 前窗口的检查发现了具体的来源伪影:正常类监控视频中不存在的编辑标题卡和平台水印。因此,视频级AUC是事件证据和事件前来源线索的组合,这种共同的区分来源会掩盖表示之间的差异。该诊断仅需要这些基准已有的标注。

英文摘要:

Articulated human pose provides detailed body-configuration information beyond coarse spatial relationships, but whether this detail yields greater discriminative information when the downstream pipeline is held fixed remains unclear. We examine this through early violence detection. Holding the tracker, temporal head, supervision, folds, and evaluation fixed, we compare five interaction representations spanning coarse bounding-box geometry, a matched handcrafted pose analogue, enriched pose descriptors, and a matched-capacity encoder learned from raw joints, under video-level evaluation with cluster-bootstrap intervals. No pose-based representation outperforms coarse geometry, though with fifteen anomalous videos this subset cannot rule out small effects. Extending the pipeline to frozen visual encoders, and repeating the comparison on XD-Violence (137 anomalous videos, nine times our UCF-Crime sample), person-crop appearance and whole-frame context both exceed geometry by a wide margin, yet context matches appearance on UCF-Crime and exceeds it on the larger split: cropping to the interacting people yields no advantage over encoding the whole frame. This prompts a direct test of what the benchmark measures. Scoring anomalous videos using only frames preceding the annotated onset, under a control removing sequence length as a cue, retains 39-91% of above-chance separation on both benchmarks, including for seven hand-designed geometric channels. Inspection of the tightest pre-onset windows identifies concrete provenance artifacts: editorial title cards and platform watermarks absent from the surveillance footage supplying the normal class. Video-level AUC here is thus a composite of event evidence and pre-event source cues, a shared source of discrimination that can obscure differences between representations. The diagnostic requires only annotations these benchmarks already ship.

↑