增强执法音频转录:基于LoRA的Whisper对随身摄像头(BWC)素材的适配
Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
浏览论文内容
中文总结 AI 辅助
针对警务音频转录的可见性悖论,该研究提出基于LoRA的Whisper适配框架,结合符号推理流程,在消费级硬件上实现高压力场景下的执法音频转录,达成93.7%词汇映射率,助力警务问责与透明度提升。
中文摘要 AI 辅助
现代警务面临“可见性悖论”:执法机构拥有PB级随身摄像头(BWC)素材,但因人工转录的人力成本过高,这些素材大多未被用于问责或系统性审查。本研究提出一种适配OpenAI Whisper架构的框架,以应对警务环境特有的声学与语言挑战。通过采用基于低秩适配(LoRA)的参数高效微调(PEFT),解决零样本模型在高压力场景、警笛声及无线电干扰下出现的显著性能下降问题。关键在于,利用8位量化与梯度检查点技术,该适配可在消费级硬件(配备NVIDIA 4GB GTX GPU的Acer Nitro本地设备)上实现。我们还将这些转录结果整合至基于领域特定本体的符号推理流程,将原始音频转化为与证据关联的事件图,实现93.7%的词汇映射率,以推进程序正义与透明度。
英文摘要
Modern policing faces a "visibility paradox" where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI Whisper architecture to the unique acoustic and linguistic challenges of the policing environment. By employing Parameter-Efficient Fine-Tuning (PEFT) through Low-Rank Adaptation (LoRA), we address the significant performance degradation observed in zero-shot models when confronted with high-stress scenarios, sirens, and radio interference. Crucially, we demonstrate that this adaptation is feasible on consumer-grade hardware (Acer Nitro local machine with NVIDIA 4GB GTX GPU) using 8-bit quantization and gradient checkpointing. We further integrate these transcriptions into a symbolic reasoning pipeline using a domain-specific ontology to transform raw audio into evidence-linked incident graphs, achieving a 93.7% lexicon mapping rate for the advancement of procedural justice and transparency.