面向临床面试培训的脚手架式多智能体大语言模型系统评估
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
浏览论文内容
中文总结 AI 辅助
本研究开发并评估了一个脚手架式多智能体大语言模型AI标准化患者培训系统,通过随机对照实验证明其能提升医学生临床访谈的沟通、共情和病史采集能力,并发布多专家标注数据集。
中文摘要 AI 辅助
临床教育必须培养医学生在不确定条件下进行安全且连贯的患者访谈的能力。传统的标准化患者(SP)培训资源密集且难以规模化。我们开发了一个脚手架式的多智能体大语言模型(LLM)AI标准化患者(AI-SP)培训平台。该系统包括一个用于模拟对话的患者智能体、一个提供苏格拉底式提示但不披露诊断信息的导师智能体,以及一个在回合级别监控临床进展但不透露总结性评分的评估智能体。在一项随机对照研究(N = 100名医学生)中,参与者被分配到多智能体(MA)脚手架条件或对照组。所有学生在指定条件下完成两次学习课程,随后在仅患者环境中进行考试。使用基于标准化客观结构化临床考试(OSCE)的评分标准评估表现。虽然两组之间的最终诊断准确性没有显著差异,但与采用结构化渐进信息披露的对照组相比,多智能体AI标准化患者系统提高了期末考试分数;最显著且一致的改进出现在沟通、共情表达和特定病史采集行为方面。这些发现表明,专门的LLM智能体提高了模拟临床访谈的过程质量,而不会人为夸大考试结果。为支持未来研究,我们发布了一个多专家标注数据集,包含转录文本、清单标注、回合级评估和OSCE对齐的评分结果。该资源旨在促进基于教学法的AI-SP系统的开发,并推进AI支持的临床推理训练研究。
英文摘要
Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.
发表机构
- The Ohio State University(俄亥俄州立大学)
- The Hong Kong University of Science and Technology(香港科技大学)
- Southern University of Science and Technology(南方科技大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Guangzhou Medical University(广州医科大学)
机构由 AI 辅助整理,请以论文原文为准。