arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

临床医生锚定的AI辅助精神科入院评估质量保障

Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

King Shi, Amanda Li, Jonathan Ivey, Synthia Qia Wang, Guan Gui, Hyunseo Kim, Peter Zandi, Jason Straub, Jacob Taylor, Ananya Joshi

arXiv 2609.21149首次发表:更新:

发表机构

Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对AI辅助精神科入院评估,提出以临床医生为中心的评估平台InterviewPlayground,通过模拟患者比较LLM与临床医生表现,发现LLM在信息恢复上更优但存在过度推断和安全问题表征不足,为质量保障奠定基础。

AI 中文摘要

在患者能够使用AI辅助精神科入院评估系统之前,卫生系统需要实用的方法,依据其临床标准对这些工具进行常规质量保障评估。由于临床医生可能采用不同的入院评估风格,针对此任务的评估必须(1)支持跨访谈方法的比较,(2)最小化临床医生的负担,以及(3)为部署这些技术的卫生系统衡量临床相关的性能。我们提出了一个以临床医生为中心的评估平台,该平台围绕一个用于开放式AI访谈的记忆增强型患者模拟器InterviewPlayground构建。我们使用InterviewPlayground结合专家撰写的病例 vignettes 创建了交互式患者,构建了一个模拟入院评估平台用于访谈,并设计了与入院评估相关的评估模式。在一项试点研究中,6名临床医生在25分钟的评估中与基于GPT的LLM入院评估访谈者进行比较,LLM从患者病例 vignettes 中恢复了更多临床相关项目(88.0%对38.9%),但做出了更多非基于访谈的临床推断(56.8%对27.8%),并且对已识别安全问题的表征频率较低(33.3%对66.7%),这为此任务的部署质量保障奠定了基础。

英文摘要

Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.

Comments7 pages, 3 figures, submitted to IAAI'27

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑