发表机构
University of Tehran; Yazd (Shahid Sadoghi) University of Medical Sciences(德黑兰大学; 亚兹德(沙希德·萨杜吉)医科大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PsyCIDRA通过双智能体框架结合自由访谈与诊断推理,在模拟和人类研究中显著提升诊断一致性,展示LLM智能体辅助精神科评估的潜力。
AI 中文摘要
大型语言模型在临床推理方面展现出潜力,但精神科访谈需要引导不断发展的对话。它们执行这种交互式评估的能力仍研究不足。我们提出了PsyCIDRA,一个双智能体框架,将自由形式的精神科访谈与诊断推理联系起来,以供专家审查。其访谈智能体使用工具来维护工作笔记、加载专家编写的技能,并检索ICD-11参考文献以指导询问。其诊断推理智能体随后接收完整的访谈记录,并报告假设以及支持、冲突和缺失的证据,当没有足够支持时,则保留最终假设。使用PsyCPG生成的患者档案,我们首先在模拟中评估PsyCIDRA。在53个评估案例中的四个模型上,它比直接提示实现了更高的诊断一致性。在81个保留的模拟案例中,排名第一的准确率为60.5%,而直接提示为51.9%。在一项包含101名人类参与者的盲法研究中,不同分组中,PsyCIDRA在79.6%的案例中与心理学家在是否提出诊断假设方面达成一致,而直接提示为65.4%。这些发现共同支持了LLM智能体通过自由形式对话协助精神科评估的潜力。通过同时检查诊断推理、访谈质量和安全性,这项研究有助于理解精神科访谈智能体的能力和局限性。
英文摘要
Large language models show promise in clinical reasoning, but psychiatric interviewing requires guiding an evolving conversation. Their ability to carry out this interactive assessment remains less studied. We present PsyCIDRA, a dual-agent framework linking free-form psychiatric interviewing with diagnostic reasoning for expert review. Its interviewer agent uses tools to maintain working notes, load expert-written skills, and retrieve ICD-11 references to guide inquiry. Its diagnostic reasoning agent then receives the completed interview transcript and reports hypotheses alongside supporting, conflicting, and missing evidence, withholding a final hypothesis when none is sufficiently supported. Using patient profiles generated with PsyCPG, we first evaluate PsyCIDRA in simulation. Across four models on 53 evaluation cases, it achieves higher diagnostic agreement than direct prompting. On 81 held-out simulated cases, rank-1 accuracy is 60.5% versus 51.9%. In a blinded study of 101 human participants in separate arms, PsyCIDRA agrees with psychologists on whether to propose a diagnostic hypothesis in 79.6% of cases, compared with 65.4% for direct prompting. Together, these findings support the potential of LLM agents to assist psychiatric assessment through free-form dialogue. By examining diagnostic reasoning, interview quality, and safety together, this study contributes to understanding the capabilities and limitations of psychiatric interview agents.
Comments41 pages, 27 figures, 20 tables; includes appendices