arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.03266cs.CLcs.AI

对设备大语言模型在临床决策支持中的基准测试与适应性研究

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

  • AI Collaborative Centre, University Health Network(人工智能协同中心,健康大学网络)
  • Princess Margaret Cancer Centre, University Health Network(彭菲癌症中心,健康大学网络)
  • Department of Electrical and Computer Engineering, University of Toronto(电气与计算机工程系,多伦多大学)
  • Division of Urology, Department of Surgery, St. Michael’s Hospital, Unity Health Toronto and University of Toronto(泌尿科,外科部,圣米歇尔医院,统一健康多伦多及多伦多大学)
  • Peter Munk Cardiac Centre, University Health Network(彼得·默克心脏中心,健康大学网络)
  • Department of Laboratory Medicine and Pathobiology and Department of Computer Science, University of Toronto(实验室医学与病理学系及计算机科学系,多伦多大学)
  • Vector Institute, Toronto(向量研究所,多伦多)

机构由 AI 辅助整理,请以论文原文为准。

Alif Munim, Jun Ma, Omar Ibrahim, Alhusain Abdalla, Shuolin Yin, Leo Chen, Bo Wang

更新

AI总结:

本文通过基准测试和微调,评估了多种设备端大语言模型在临床任务中的性能,发现其在诊断准确率上可与开源模型和专有模型相媲美,并展示了微调对提升诊断准确性的效果。

AI中文摘要:

大型语言模型(LLMs)在临床决策中迅速发展,但专有系统的部署受到隐私问题和对云基础设施依赖的阻碍。开源替代方案允许本地推断,但通常具有较大的模型规模,限制了其在资源有限的临床环境中的使用。本文对gpt-oss(20b,120b)、Qwen3.5(9B,27B,35B)和Gemma 4(31B)等设备端LLM在三个代表性临床任务上进行基准测试:一般疾病诊断、专科诊断与管理以及人类专家评分和评估。我们比较了这些模型与最先进的专有模型(GPT-5.1、GPT-5-mini和Gemini 3.1 Pro)和领先的开源模型(DeepSeek-R1)的性能,并进一步通过在一般诊断数据上微调gpt-oss-20b和Qwen3.5-35B来评估设备端系统的适应性。在各项任务中,设备端模型在性能上可与或超过DeepSeek-R1和GPT-5-mini,尽管模型规模显著较小。此外,微调显著提高了诊断准确性,微调后的Qwen3.5-35B达到87.9%,接近专有GPT-5.1(89.4%)。在基础设备端模型中,Gemma 4 31B在一般诊断准确率上表现最强,达到86.5%,超过GPT-5-mini并接近微调后的Qwen3.5-35B。错误分析揭示了所有模型中87.2%的诊断错误是临床合理的鉴别诊断而非离题预测,上界分析显示通过改进答案选择可达到93.2%的准确率。这些发现突显了设备端LLM在提供准确、适应性强且隐私保护的临床决策支持方面的潜力,为LLM在日常临床实践中的广泛应用提供了可行路径。

英文摘要:

Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy concerns and reliance on cloud-based infrastructure. Open-source alternatives allow local inference but often have large model sizes that limit their use in resource-constrained clinical settings. Here, we benchmark on-device LLMs from the gpt-oss (20b, 120b), Qwen3.5 (9B, 27B, 35B), and Gemma 4 (31B) families across three representative clinical tasks: general disease diagnosis, specialty-specific (ophthalmology) diagnosis and management, and simulation of human expert grading and evaluation. We compare their performance with state-of-the-art proprietary models (GPT-5.1, GPT-5-mini, and Gemini 3.1 Pro) and a leading open-source model (DeepSeek-R1), and we further evaluate the adaptability of on-device systems by fine-tuning gpt-oss-20b and Qwen3.5-35B on general diagnostic data. Across tasks, on-device models achieve performance comparable to or exceeding DeepSeek-R1 and GPT-5-mini despite being substantially smaller. In addition, fine-tuning remarkably improves diagnostic accuracy, with the fine-tuned Qwen3.5-35B reaching 87.9% and approaching the proprietary GPT-5.1 (89.4%). Among base on-device models, Gemma 4 31B achieved the strongest general diagnostic accuracy at 86.5%, exceeding GPT-5-mini and approaching the fine-tuned Qwen3.5-35B. Error characterization revealed that 87.2% of diagnostic errors across all models were clinically plausible differentials rather than off-topic predictions, and upper-bound analysis showed up to 93.2% attainable accuracy through improved answer selection. These findings highlight the potential of on-device LLMs to deliver accurate, adaptable, and privacy-preserving clinical decision support, offering a practical pathway for broader integration of LLMs into routine clinical practice.

↑