理解使用大型语言模型的临床认知对话
Understanding Clinical Cognitive Dialogues Using Large Language Models
- College of AI, Cyber and Computing(人工智能、网络与计算学院)
- University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究构建了33次认知评估对话的标注语料库,并基准测试大型语言模型,发现指令调优和推理感知微调分别提升生成与分类性能,但细粒度对话行为区分仍具挑战性。
AI中文摘要:
面对面的认知评估既是一次测试,也是一次互动。临床医生解释任务、修复误解并适应患者的反应,而患者可能会犹豫、寻求澄清或脱离互动。然而,临床对话资源很少标注研究这些行为所需的大规模互动结构。我们提供了一个去标识化的语料库,包含33次认知评估对话,共8,250个话语,标注了三种说话者角色和56种对话行为。我们使用该语料库对大型语言模型进行细粒度对话行为分类和下一患者话语生成的基准测试。我们还测试了域外指令数据和解释增强训练是否迁移到这一临床环境。指令调优产生了最强的患者话语参考匹配,并提高了分类准确性。推理感知微调在LLaMA-3.1-8B变体中产生了最强的分类结果。然而,即使是最好的模型也难以区分密切相关的对话行为,这表明广泛的对话意图比细粒度的交际功能更容易识别。该语料库和基准使认知评估中的互动结构可测量,并支持关于对话标记、临床医生教育和经过仔细验证的模拟患者的后续工作。这项工作不做出诊断声明。相反,它提供了研究这些应用所需的数据和评估框架。
英文摘要:
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.