基于特征的法庭科学似然比:将神经网络与贝叶斯概率计算相结合
Feature-Based Likelihood Ratios for Forensic Science: Combining Neural Networks with Bayesian Probability Calculus
浏览论文内容
中文总结 AI 辅助
本文提出结合神经网络与贝叶斯概率计算的梯度下降方法,训练两级模型以获得校准良好的基于特征的似然比,在玻璃碎片数据集上较现有系统平均提升4.5倍,但仍存在1.5倍差距。
中文摘要 AI 辅助
在法庭科学中,当存在犯罪现场证据(CSE)和与嫌疑人相关的证据(SRE)时,通常以似然比(LR)的形式报告该证据的价值。LR 可以计算为给定 SRE 时 CSE 的概率除以给定 CSE 由从替代罪犯群体中随机选择的人生成时的概率。在法庭科学中,这被称为基于特征的 LR,直观上 LR 通过“典型性”对比“相似性”。由于为数据找到合适的模型通常是基于特征的 LR 的一个问题,人们要么求助于基于评分的 LR,要么事后将基于特征的输出调整为校准良好的输出。无论哪种方式,上述 LR 的定义都被破坏,LR 作为 CSE 与 SRE 之间的相似性除以 CSE 的典型性的解释也被破坏。在这里,我们报告了在获得即时表现良好的基于特征的 LR 方面的进展,使用梯度下降结合贝叶斯概率理论来训练一个两级模型,这是法庭科学中描述连续数据分布的主要 LR 模型。对于来自法庭案例工作的玻璃碎片上的激光剥蚀电感耦合等离子体质谱测量数据集,我们表明,在验证数据上我们最好的模型在测试集上产生的基于特征的 LR 的校准程度远优于在同一类型数据上训练的最先进的基于特征的 LR 系统,并且在 $C_{\mathrm{llr}}$ 上平均提高了 4.5 倍。对于以“相似性”对比“典型性”来解释的 LR,这是一个重大进步。然而,对于这些数据,最先进的 LR 系统在 $C_{\mathrm{llr}}$ 上仍然表现好 1.5 倍。我们还提出了未来计划以缩小这一剩余差距。为了促进合作,我们已将相关代码放在 GitHub 上。
英文摘要
In forensic science, when crime-scene evidence (CSE) and suspect-related evidence (SRE) is present, it is customary to report on the value of this evidence in the form of a likelihood ratio (LR). The LR can be calculated as the probability of CSE given SRE divided by the probability of CSE given that it was generated by a randomly selected person from an alternative culprit population. In forensic science, this is known as a feature-based LR, and intuitively the LR contrasts "similarity" by "typicality". Since it is generally a problem for feature-based LRs to find appropriate models for the data, one either resorts to score-based LRs or to adjusting the feature-based output post-hoc to well-calibrated output. Either way, the above definition of the LR is broken and interpretation of the LR as similarity between CSE and SRE divided by typicality of CSE is destroyed. Here, we report on progress in obtaining instantly well-performing feature-based LRs using gradient descent in combination with Bayesian probability theory to train a two-level model, the main LR model in forensic science for describing distributions of continuous data. For a dataset of laser-ablation inductively-coupled-plasma mass-spectrometry measurements on glass fragments from forensic casework, we show that our best model on validation data yields much better calibrated feature-based LRs on the test set when compared to state-of-the-art feature-based LR systems trained on the same type of data, and that it improves a factor of 4.5 on average on $C_{\mathrm{llr}}$. For LRs interpretable in terms of "similarity" contrasting "typicality", this is a major advancement. However, a state-of-the-art LR system still performs a factor of 1.5 better on $C_{\mathrm{llr}}$ for this data. We also present future plans to close this remaining gap. In order to facilitate collaboration, we have put relevant code on GitHub.
发表机构
- Netherlands Forensic Institute(荷兰法医研究所)
机构由 AI 辅助整理,请以论文原文为准。