发表机构
Research Center for Social Computing and Interactive Robotics; Harbin Institute of Technology(社会计算与信息检索研究中心; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过令牌匹配实验发现,临床数据主要提升临床任务且兼顾知识任务,说教性数据仅提升知识任务,并指出医学大语言模型的数据配比应以应用为导向,推理场景应增加临床数据比例。
AI 中文摘要
医学大语言模型通常使用说教性数据(如教科书)和临床数据(如患者记录)的混合数据进行训练,然而这些数据类型如何差异化地塑造模型能力仍不清楚。我们通过令牌匹配实验来解决这一问题,实验中变化说教性数据与临床数据的比例,并分析数据组成如何影响模型在知识密集型任务和临床导向任务上的性能、能力画像及错误模式。我们发现了跨任务类型的不对称迁移:临床数据能提升临床导向任务,同时在知识密集型任务上保持竞争力,而说教性数据主要提升知识密集型任务。错误分析提示存在“知行差距”,即知识回忆的提升并不能可靠地泛化到临床推理中。我们进一步观察到,适量的临床数据即可在基于电子健康记录(EHR)的任务上带来大部分收益,而最优混合比例会随下游任务的知识需求和临床推理需求而变化。这些发现表明,医学大语言模型的数据策展应以应用为导向,对于推理密集型使用场景,应优先采用更高比例的临床数据。
英文摘要
Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition affects performance, capability profiles, and error patterns across knowledge-intensive and clinic-oriented tasks. We uncover an asymmetric transfer across task types: clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones, whereas didactic data mainly improves knowledge-intensive tasks. Error analysis suggests a knowing-doing gap, where improvements in knowledge recall do not reliably generalize to clinical reasoning. We further observe that modest amounts of clinical data yield most of the gains on EHR-grounded tasks, while the optimal mixture ratio varies with the knowledge and clinical reasoning demands of downstream tasks. These findings suggest that medical LLM data curation should be application-driven, with higher proportions of clinical data preferred for reasoning-intensive use cases.
CommentsThis paper is accepted by NLPCC 2026 oral