发表机构
HelpMum(HelpMum)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估了针对尼日利亚孕产妇和疫苗接种领域微调的MamaBot-Llama和Vax-Llama与基础模型Llama-3.1-8B-Instruct,发现基于高质量数据微调可提升性能,但数据质量不足时性能下降,强调临床部署前需严格验证。
AI 中文摘要
背景:大型语言模型(LLM)可改善低资源环境下的医疗信息传递,但可能产生不准确或文化上不合适的建议。本研究评估了针对尼日利亚孕产妇健康和疫苗接种的领域特定微调。目标:比较HelpMum的MamaBot-Llama和Vax-Llama与Meta的Llama-3.1-8B-Instruct在准确性、安全性、清晰度、情境适当性和可信度方面的表现。方法:我们评估了200个医疗问题,孕产妇健康和疫苗接种各100个,覆盖每个领域的五个子领域。MamaBot-Llama和Vax-Llama分别使用低秩适配(LoRA)在超过36,000个孕产妇健康和9,000个疫苗接种问答对上进行微调。两名尼日利亚执业医师独立使用5点李克特量表对回答进行评分。配对比较采用Wilcoxon符号秩检验。结果:性能因领域而异。MamaBot-Llama在所有标准上显著优于基础模型,总体提升4.9%(p<.001),包括临床可信度提升(+7%)和医学准确性提升(+5%)。关键问题减少50%,临床医生在78%的案例中更偏好它。相比之下,Vax-Llama总体下降5.2%(p<.001),关键问题增加192%,安全问题增加400%。结论:基于高质量、临床医生策划的数据进行领域特定微调可以改善医疗LLM性能,但当数据集质量不足时也可能降低性能。在临床部署前,必须进行严格的领域特定验证。医师评估者提供了知情同意,聊天机器人日志已匿名化。关键词:大型语言模型;微调;孕产妇健康;疫苗接种;医疗人工智能;低资源环境;尼日利亚;模型评估;LoRA;医学准确性
英文摘要
Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally inappropriate advice. This study evaluated domain-specific fine-tuning for maternal health and vaccination in Nigeria. Objective: To compare HelpMum's MamaBot-Llama and Vax-Llama with Meta's Llama-3.1-8B-Instruct for accuracy, safety, clarity, contextual appropriateness, and trustworthiness. Methods: We evaluated 200 healthcare questions, 100 each for maternal health and vaccination, across five subdomains per domain. MamaBot-Llama and Vax-Llama were fine-tuned using Low-Rank Adaptation on over 36,000 maternal health and 9,000 vaccination question-answer pairs, respectively. Two Nigerian licensed physicians independently rated responses using a 5-point Likert scale. Paired comparisons used Wilcoxon signed-rank tests. Results: Performance varied by domain. MamaBot-Llama significantly outperformed the base model across all criteria, with a 4.9% overall improvement (p < .001), including gains in clinical trustworthiness (+7%) and medical accuracy (+5%). Critical issues decreased by 50%, and clinicians preferred it in 78% of cases. In contrast, Vax-Llama showed a 5.2% overall decline (p < .001), with critical issues increasing by 192% and safety concerns by 400%. Conclusions: Domain-specific fine-tuning can improve healthcare LLM performance when based on high-quality, clinician-curated data, but may also degrade performance when dataset quality is inadequate. Rigorous domain-specific validation is essential before clinical deployment. Physician evaluators provided informed consent, and chatbot logs were anonymized. Keywords: Large language models; Fine-tuning; Maternal health; Vaccination; Healthcare AI; Low-resource settings; Nigeria; Model evaluation; LoRA; Medical accuracy