我们能否预判暴力?基于事件前行为线索的多模态学习
Can We Anticipate Violence? Multimodal Learning from Pre-Incident Behavioral Cues
浏览论文内容
中文总结 AI 辅助
本研究利用多模态视频信号(面部外观、音频和身体运动)构建短时域事件前风险识别模型,通过XD-Violence数据集验证,最佳配置DeiT-Tiny组合所有模态达到91.21%准确率,证明互补线索可有效预判暴力风险。
中文摘要 AI 辅助
在暴力行为开始后检测暴力固然重要,但识别其发生前即刻出现的行为线索同样关键。本研究探讨了从多模态视频信号中进行短时域事件前风险识别的问题。我们利用带时间标注的XD-Violence片段构建了一个二分类的正常与风险场景,使用443个样本,并在训练集、验证集和测试集之间进行了源级分离。每个样本包含一个长度可变的事件前片段,其时长由事件前可观察的行为背景决定。事件本身被排除在所有输入片段之外。我们评估了三种互补的信息来源:面部区域外观、时间对齐的音频以及从跟踪关键点导出的身体运动特征。在相同的划分下,使用Swin-Tiny、ViT-Tiny和DeiT-Tiny进行了受控消融实验,以衡量每种模态的贡献。结果表明,组合所有模态比单独使用任何其他组合更有效。最佳配置为DeiT-Tiny结合音频、面部外观和运动,在留出测试集上达到了91.21%的准确率、88.96%的平衡准确率、93.65%的F1分数和96.38%的ROC-AUC。这些结果表明,互补的外观、声学和运动学线索为识别升高的事件前风险提供了有用的证据。
英文摘要
Detecting violence after it begins is important from recognizing behavioral cues that appear immediately beforehand. This work studies short-horizon pre-incident risk recognition from multimodal video signals. We construct a binary Normal-versus-Risky setting from temporally annotated XD-Violence clips, using 443 samples with source-level separation across training, validation, and test sets. Each sample consists of a variable-length pre-incident clip, with its duration determined by the observable behavioral context preceding the incident. The inci- dent itself is excluded from all input clips. We evaluate three complementary information sources: facial-region appearance, temporally aligned audio, and body-motion features derived from tracked keypoints. Controlled ablations are performed with Swin-Tiny, ViT-Tiny, and DeiT-Tiny to measure the contribution of each modality under the same split. Results show that combining all modalities is more effective than using any other combination alone. The best configuration, Deit-Tiny with audio, facial appearance, and motion, achieves 91.21% accuracy, 88.96% balanced accuracy, 93.65% F1-score, and 96.38% ROC-AUC on the held-out test set. These results suggest that complementary appearance, acoustic, and kinematic cues provide useful evidence for recognizing elevated pre-incident risk.
发表机构
- The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。