arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22460cs.LGcs.CY

MASH-Bench:大规模枪击风险分类中的跨源故障诊断

MASH-Bench: Diagnosing Cross-Source Failure in Mass-Shooting Risk Classification

  • Virginia Commonwealth University(弗吉尼亚联邦大学)

机构由 AI 辅助整理,请以论文原文为准。

Neha Sharma, Ritesh Sharma

中文总结 AI 辅助

本文提出MASH-Bench基准,通过实验发现大规模枪击风险分类的跨源泛化受特征完整性和标签流行度限制,DANN可部分提升GVA的极高风险召回率,为相关研究提供受控环境。

中文摘要 AI 辅助

公开的大规模枪击数据库在覆盖范围、特征可用性和报告实践方面存在显著差异,给需跨数据源泛化的机器学习模型带来挑战。本文提出MASH-Bench,这是一个由美国四个数据库(Kaggle、Mother Jones、Stanford MSA和枪支暴力档案GVA)中6968起事件构成的统一基准。采用留一数据集交叉验证(LODO)评估跨源风险分类性能:随机森林、XGBoost和LightGBM在整理后的数据源上的极高风险(VeryHigh-risk)召回率为0.68-0.89,但向GVA泛化时表现极差,平均召回率降至0.20,精确率仅为0.0004。为探究性能下降的原因,本文开展受控特征掩码消融实验,从整理后的数据源中移除GVA缺失的5个特征,结果召回率骤降至0,表明特征完整性是导致跨源故障的主要因素。进一步评估三种领域适应方法:DANN使GVA上的极高风险召回率提升0.282(95%置信区间[0.11,0.47],p=0.003),但精确率仍低;CORAL和重要性加权的召回率为0;Oracle先验偏移校准也无法恢复极高风险预测,说明仅修正标签侧在观测到的特征缺陷下不足。按组审计还发现与媒体归因的心理健康标签相关的显著差异。总体而言,MASH-Bench的结果表明,跨源泛化更多受特征完整性和标签流行度限制,而非分类器选择;该基准为诊断跨源风险分类中的这些效应提供了受控环境。

英文摘要

Public mass-shooting databases differ substantially in coverage, feature availability, and reporting practices, creating challenges for machine-learning models that must generalize across data sources. We introduce MASH-Bench, a harmonized benchmark of 6,968 incidents from four U.S. databases: Kaggle, Mother Jones, Stanford MSA, and the Gun Violence Archive (GVA). We evaluate cross-source risk classification using leave-one-dataset-out (LODO) evaluation. Random Forest, XGBoost, and LightGBM achieve VeryHigh-risk recall of 0.68-0.89 on the curated sources but generalize poorly to GVA, where mean recall drops to 0.20 and precision to 0.0004. To investigate the source of this degradation, we conduct a controlled feature-masking ablation that removes the five features unavailable in GVA from the curated sources. The resulting recall collapse to zero provides evidence that feature completeness is a major contributor to the observed cross-source failure. We further evaluate three domain-adaptation approaches: DANN, CORAL, and importance weighting. DANN improves VeryHigh-risk recall on GVA by 0.282 (95% CI [0.11, 0.47], p = 0.003), although precision remains low, whereas CORAL and importance weighting yield zero recall. Oracle prior-shift recalibration likewise fails to recover VeryHigh-risk predictions, indicating that label-side correction alone is insufficient under the observed feature deficiencies. A per-group audit further identifies substantial disparities associated with media-attributed mental-health labels. Overall, these results indicate that, in MASH-Bench, cross-source generalization is constrained more by feature completeness and label prevalence than by classifier choice. The benchmark provides a controlled setting for diagnosing these effects in cross-source risk classification.

↑