arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为临床预测统一多模态学习者的大语言模型

Large Language Models as Unified Multimodal Learners for Clinical Prediction

Ajay Madhavan Ravichandran, Bilgin Osmandoja, Klemens Budde, Klaus Netter, Tobias Strapatsas, Aljoscha Burchardt, Sebastian Möller, Roland Roller

arXiv 2607.15380首次发表:更新:

发表机构

German Research Center for Artificial Intelligence (DFKI); Charité Universitätsmedizin Berlin; DNC Information Management GmbH; Klinik für Akut- und Notfallmedizin, Asklepios Klinikum Harburg; Technical University Berlin(德国人工智能研究中心(DFKI); 柏林夏里特大学医学中心; DNC信息管理有限公司; 哈堡阿斯克勒庇俄斯医院急性与急诊医学科; 柏林工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨利用大语言模型进行临床预测,提出将不同模态患者数据转为自然语言序列,端到端微调预训练语言模型的方法,经三个临床任务评估,该方法匹配或超越多模态基线,优于临床梯度提升系统,降低系统复杂性。

AI 中文摘要

电子健康记录将自由文本临床叙述与结构化测量相结合。然而,大多数临床预测系统仍依赖特定任务融合架构。本文提出更简单的方法:将所有患者数据转换为单一自然语言序列,对预训练语言模型进行端到端微调,无需架构修改用于融合。在三个临床预测任务中评估该方法,结果表明统一文本序列化匹配或超越特定任务多模态基线,在移植失败预测上优于临床部署的梯度提升系统。这表明单一基于序列化的范式足以进行多模态临床预测,可大幅降低系统复杂性。

英文摘要

Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities. Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mechanisms that must be re-engineered for every new task and clinical setting. We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language model end-to-end, with no architectural modification for fusion. We evaluate this approach across three clinically distinct prediction tasks: in-hospital mortality on MIMIC-III, graft failure prediction using longitudinal data from a German transplant center, and emergency triage classification from ambulance records - comparing encoder-based (ModernBERT) and decoder-based (Llama 3.1, Gemma, DeepSeek-R1-Qwen, Qwen3) fine-tuning against established multimodal baselines and, for graft failure, a gradient boosting model currently used in clinical practice for post-transplant patient management. Across all three tasks, unified textual serialization matches or exceeds task-specific multimodal baselines, and outperforms the clinically deployed gradient boosting system on graft failure prediction. These results indicate that a single serialization-based paradigm, without bespoke fusion architectures, is sufficient for multimodal clinical prediction - substantially reducing system complexity while matching or exceeding specialized designs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑