arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向精神病学领域的专用大语言模型在回答患者问题时的性能表现

Performance of a domain-specific large language model in answering patient questions in psychiatry

Alexander J. Hish, Arjun Nagendran, Scott N. Compton

arXiv 2608.22797首次发表:更新:

AI 中文总结

本研究开发了专用LLM MIND,对比ChatGPT、OpenEvidence回答患者关于艾司西酞普兰的问题,发现MIND评分更高但医生更偏好ChatGPT,为构建精神病学安全LLM提供了进展。

AI 中文摘要

背景:本研究旨在评估仅基于患者教育资源训练的专用大语言模型(LLM)是否能以优于通用LLM聊天机器人的方式回答关于精神科药物的问题。我们开发了一款针对临床准确性进行微调的LLM(名为“MIND”),该模型基于权威医疗机构的患者教育资源训练。方法:我们通过两种方法比较了MIND、ChatGPT和OpenEvidence对关于艾司西酞普兰(escitalopram)的患者问题的回答:(1)根据包含准确性、清晰度、完整性、细致度、安全性及转诊适宜性的评分标准进行计算机分析;(2)由10名持有执照的精神科医生对类似指标进行评分。结果:根据评分标准,MIND在所有维度的评分均最高(p<0.001);由精神科医生评分时,ChatGPT被评为准确的频率略高于MIND,效应量可忽略(p=0.021,r=0.073),MIND被评为完整的频率高于ChatGPT,效应量较小(p<0.001,r=0.160),且两者被评为安全的频率相同(p=0.955,r=0.002);多数精神科医生更偏好ChatGPT生成的回答(57.6%),而非MIND(42.4%,p=0.003)。结论:MIND多数情况下能以精神科医生认为准确、完整且安全的方式回答许多关于艾司西酞普兰的问题,尽管其回答更完整,但精神科医生仍更偏好ChatGPT的回答,MIND为构建用于增强精神病学患者教育的安全LLM系统迈出了一步。

英文摘要

Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM ("MIND") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑