评估阿拉伯语危机求助电话中的自杀风险:阿拉伯语与英语大语言模型的对比
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
浏览论文内容
中文总结 AI 辅助
本研究对比阿拉伯语与英语大语言模型,利用黎巴嫩匿名危机求助热线转录本评估自杀风险,发现两类模型均可有效分类,英语模型表现更优,极高危案例区分度更高。
中文摘要 AI 辅助
危机求助热线通过结构化访谈评估自杀风险,该过程耗时且依赖接线员的培训与工作量。自然语言处理可支持风险评估与电话优先级排序,但几乎没有研究针对阿拉伯语求助热线电话开展,也未考虑真实热线数据的隐私约束。我们分析了黎巴嫩国家情感支持与自杀预防生命线的匿名通话记录:音频从未离开热线,通话通过黎凡特阿拉伯语语音识别模型现场转录,阿拉伯语命名实体识别模型在本地移除了可识别信息,仅将匿名转录本分享给研究团队。接线员记录了哥伦比亚自杀严重程度量表的5项自杀意念条目,我们将其合并为两个二元结局:高危与极高危。我们还将转录本机器翻译为英语,形成阿拉伯语/英语配对对比。在每个语料库上,我们对5个指令调优大语言模型及6个Transformer编码器基线模型(4个阿拉伯语、2个英语)进行微调,并在保留的测试集上评估所有模型。共纳入383通电话:373通用于高危任务(52.3%为阳性),297通用于极高危任务(30.0%为阳性)。最佳阿拉伯语模型在极高危任务上达到宏F1值81.19、ROC-AUC值90.61;最佳英语模型达到85.00、92.59,识别出88.9%的极高危电话。两种语言中,极高危电话的区分度均高于高危电话,翻译为英语未降低最佳观测性能。研究表明,可在不将音频传出热线的情况下,从匿名阿拉伯语转录本中对自杀风险进行分类;极高危任务的结果支持其作为面向接线员的工具开展进一步测试,而低严重程度意念则是更难的案例。
英文摘要
Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresses Arabic-language helpline calls or operates within the privacy constraints of real helpline data. We analysed de-identified transcripts from Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Audio never left the helpline: calls were transcribed on site with a speech recognition model for Levantine Arabic, and an Arabic named-entity recognition model removed identifying information locally. Only the de-identified transcripts were shared with the research team. Operators recorded the five suicidal ideation items of the Columbia Suicide Severity Rating Scale, which we combined into two binary outcomes: at-risk and high-risk. We also machine-translated the transcripts into English, giving a paired Arabic/English comparison. On each corpus, we fine-tuned five instruction-tuned large language models alongside six transformer encoder baselines (four Arabic, two English) and evaluated all models on a held-out test set. We included 383 calls: 373 for the at-risk task (52.3% positive) and 297 for the high-risk task (30.0% positive). The best Arabic model reached a macro-F1 of 81.19 and a ROC-AUC of 90.61 on high-risk; the best English model reached 85.00 and 92.59, identifying 88.9% of high-risk calls. In both languages, high-risk calls separated more cleanly than at-risk calls, and translation to English did not reduce the best observed performance. Suicide risk can be classified from de-identified Arabic transcripts without sending audio outside the helpline. The high-risk results support further testing as an operator-facing tool; lower-severity ideation proved the harder case.
发表机构
- Yale University(耶鲁大学)
- American University of Beirut(贝鲁特美国大学)
- Embrace, Mental Health Center(Embrace心理健康中心)
机构由 AI 辅助整理,请以论文原文为准。