发表机构
ESRC Vulnerability and Policing Futures Research Centre; School of Law, University of Leeds(经济与社会研究理事会脆弱性与警务未来研究中心; 利兹大学法学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨基于美国开源警方数据的语言模型分类管道能否用于估计英国警方事件叙述中四种脆弱性指标流行率。通过多阶段管道分析近3000份日志,虽能产生流行率估计,但单纯部署不可靠,纠正偏差需大量人工和统计调整,强调了该方法的潜力与局限。
AI 中文摘要
目的:了解日常警务中涉及弱势群体的程度可为资源配置、培训和多机构应对提供信息,但行政数据提供的见解有限。我们探讨基于美国开源警方数据开发的基于语言模型的分类管道是否可用于估计英国警方事件叙述中四种脆弱性指标(心理健康问题、药物滥用、酒精依赖和无家可归)的流行率,以及何时输出可被视为合理的测量值。方法:我们使用结合重复模型推理、标签聚合、结构化人工审查和统计校正的多阶段管道分析了来自英国一支警察部队的近3000份去识别化事件日志。该管道在本地托管的开放权重语言模型上运行,反映了警方必须工作的安全环境。结果:语言模型可以大规模产生有意义的(即使不完美)流行率估计。心理健康问题指标在大约五分之一的事件中出现,其他指标的流行率较低。然而,单纯的语言模型部署是不可靠的:单次分类不稳定,聚合输出相对于人工判断系统地过度分配指标。纠正这些偏差需要大量的人工投入和统计调整,仍存在相当大的不确定性。结论:虽然语言模型可以从未结构化的警方数据中提取信息,但在没有仔细的方法支持的情况下,其输出不能被视为有效的测量值。在总体层面,可实现合理的估计,但资源密集;在个体层面,错误仍然频繁且不可预测,限制了其在操作决策中的适用性。本研究强调了基于语言模型的测量在应用环境中的潜力和局限性。
英文摘要
Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based classification pipeline, developed on open-source US police data, can be adapted to estimate the prevalence of four vulnerability indicators - mental ill health, substance misuse, alcohol dependence, and homelessness - in UK police incident narratives, and when outputs can be treated as defensible measurements. Methods: We analyse nearly 3,000 de-identified incident logs from a UK police force, using a multi-stage pipeline combining repeated model inference, label aggregation, structured human review, and statistical correction. The pipeline runs on a locally hosted open-weight LLM, reflecting the secure environments police must work in. Results: LLMs can produce meaningful, if imperfect, prevalence estimates at scale. Mental ill health indicators are present in approximately one in five incidents, with lower prevalence for other indicators. However, naive LLM deployment is unreliable: single-pass classifications are unstable, and aggregated outputs systematically over-assign indicators relative to human judgement. Correcting these biases required substantial human input and statistical adjustment, leaving considerable uncertainty. Conclusions: While LLMs can extract information from unstructured police data, their outputs cannot be treated as valid measurements without careful methodological support. At the population level, defensible estimates are achievable but resource-intensive; at the individual level, errors remain frequent and unpredictable, limiting suitability for operational decisions. This study highlights both the potential and the constraints of LLM-based measurement in applied settings.
Comments25 pages, 4 figures. Preprint. v2: revised following peer review