可本地部署的小型语言模型用于急诊科决策支持:微调策略的系统基准测试
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
浏览论文内容
中文总结 AI 辅助
该研究针对急诊科决策支持的隐私风险,通过基准测试发现经LoRA微调的开源小型语言模型在分诊和转诊任务上优于商业模型,可本地部署且具备临床竞争力。
中文摘要 AI 辅助
在急诊科(ED)部署大型语言模型(LLM)用于决策支持面临两大挑战:将患者数据传输给闭源商业LLM带来的隐私风险,以及针对可本地部署的开源小型语言模型(SLM)的微调策略缺乏系统评估。我们使用零样本提示、前缀调优、低秩适配(LoRA)和全量微调,在三项急诊科任务(分诊级别预测、专科转诊推荐和诊断预测)上对8个开源SLM进行基准测试。使用2083个MIMIC-IV-ED病例,以Claude Haiku 4.5和Claude Sonnet 4.5作为基准,我们发现经LoRA微调的开源SLM在分诊级别预测和专科转诊推荐上优于商业基准,而诊断预测对开源SLM仍具挑战性。混淆矩阵分析进一步显示,经微调的开源SLM能够检测出商业基准遗漏的最高严重程度患者。这些结果表明,可本地部署的SLM可在急诊科决策支持中达到临床可比的性能。
英文摘要
Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting patient data to closed-source commercial LLMs and the lack of systematic evaluation of fine-tuning strategies for locally deployable open-source small language models (SLMs). We benchmarked eight open-source SLMs using zero-shot prompting, prefix tuning, Low-Rank Adaptation (LoRA), and full fine-tuning on three ED tasks: triage level prediction, specialist referral recommendation, and diagnosis prediction. Using 2,083 MIMIC-IV-ED cases and Claude Haiku 4.5 and Claude Sonnet 4.5 as baselines, we found that LoRA fine-tuned open-source SLMs outperform commercial baselines on triage level prediction and specialist referral recommendation, while diagnosis prediction remains challenging for open-source SLMs. Confusion matrix analysis further shows that fine-tuned open-source SLMs can detect highest-severity patients missed by the commercial baselines. These results demonstrate that locally deployable SLMs can achieve clinically competitive performance for ED decision support.