发表机构
Heidelberg University; University of Oxford(海德堡大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过分析HELP-Med数据集的1800条对话,评估GPT 4o、Llama 3和Command R+三种LLM的医疗领域社会交际能力,发现其结构化表现差,需调整评估框架以适配人机差异。
AI 中文摘要
背景:有效的临床实践高度依赖医护人员的社会交际技能,大型语言模型(LLMs)已被提出用于患者分诊、报告起草或医学术语翻译等任务,以支持知情决策,这些应用既需要事实能力也需要社会能力。本研究评估LLMs与参与者的对话,以评估LLMs生成文本中展现的社会交际能力的当前状态。方法:我们从HELP-Med数据集中提取了一部分扩展对话,包含1800条人类参与者寻求医疗信息与三种不同LLMs(GPT 4o、Llama 3和Command R+)互动的对话记录,两名专家使用IC-MD工具(最初用于评估医学生招生中的互动能力)对记录进行编码,以分析社会交际行为(非敌意、敏感性、结构化、非侵入性)的表现。结果:研究中的LLMs在非敌意方面表现出色,在敏感性和非侵入性方面结果不一,在结构化方面表现较差。结论:当前的LLMs缺乏作为医疗顾问安全有效使用所需的一致且可靠的社会交际技能,尽管现有的互动能力评估框架可能支持开发更具社会响应性的LLMs,但这些框架需要调整,以考虑人类与LLMs之间理想行为的差异。
英文摘要
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.