arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17270cs.CLcs.CY

不可转移的安全性:可部署医学语言模型中的跨语言临床正确性漂移

Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models

Anthonio Oladimeji Gabriel, Dimeji Olawuyi, Toba Ajayi, Temilola Aderemi

首次发表
浏览论文内容

中文总结 AI 辅助

研究资源匮乏健康环境中可部署医学语言模型的跨语言临床正确性漂移,构建英语-豪萨语问题对评估六个模型,发现本地可部署模型临床正确性下降,前沿模型较稳定,缺陷源于可部署层而非语言或临床材料。

中文摘要 AI 辅助

大语言模型的安全性评估主要在英语环境和前沿系统中进行。本文研究在资源匮乏的健康环境中,用当地语言运行的小型量化系统的情况。构建了英语-豪萨语匹配问题对,针对尼日利亚北部疟疾、镰状细胞病和结核病三种高负担疾病情况,对六个模型进行评估。结果显示,本地可部署模型的平均临床正确性从英语的1.57降至豪萨语的-0.03,前沿模型从2.00降至1.75且无有害回答。漂移在所有三种情况中一致,评分者间临床正确性一致性高,危害判断一致性低。缺陷是可部署层的特性,而非语言或临床材料的问题。

英文摘要

Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small quantised systems are run locally and queried in local languages. We ask whether clinical safety established in English transfers to Hausa, and whether any failure is attributable to the language, the clinical task, or the class of model that low-resource deployment admits. Matched English-Hausa question pairs were built for three conditions of high burden in northern Nigeria: malaria, sickle cell disease, and tuberculosis, probing knowledge recall, emergency triage, a leading question inviting a contraindicated action, and a traditional-remedy claim. Six models were evaluated: five locally deployable systems of 4-9 billion parameters, two medically fine-tuned, and one frontier system. All 128 responses were scored against Nigerian national treatment guidelines by two fluent Hausa speakers working independently and blind to one another. Among locally deployable models, mean clinical correctness fell from 1.57 in English to -0.03 in Hausa, on a scale where 2 denotes a correct answer and -1 an actively harmful one. The frontier model moved from 2.00 to 1.75 and produced no response judged harmful in either language. Drift was consistent across all three conditions. Inter-rater agreement was substantial for clinical correctness (kappa = 0.70); agreement on harm was initially poor (kappa = 0.22) and is examined in detail. Because a frontier model answers the same questions competently in Hausa, the deficit is a property neither of the language nor of the clinical material, but of the deployable tier.

发表机构

  • AI SafetyX(人工智能安全X)

机构由 AI 辅助整理,请以论文原文为准。

↑