arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20453cs.CLcs.AIcs.LG

一种用于将大语言模型零样本适应谵妄预测的知识注入框架

A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction

Jessica Sena, Shesadree Priyadarshani, Miguel Contreras, Bharat Gandhi, Scott Siegel, Subhash Nerella, Parisa Rashidi

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型零样本适应谵妄预测中领域知识不足问题,提出轻量级知识注入框架。在推理时用外部临床知识报告增强数据摘要,无需微调或检索。实验表明该方法能提升模型性能,缩小与前沿模型差距,且收益依赖临床意义内容。

中文摘要 AI 辅助

大语言模型在临床预测方面有潜力,但在专业任务上的零样本性能受限于不完整的领域知识,尤其是对于较小的可本地部署模型。我们提出了一个轻量级知识注入框架,用于零样本ICU谵妄预测。在推理时,该框架用外部临床知识报告增强结构化电子健康记录数据的确定性自然语言摘要,无需微调或检索。我们在MIMIC IV数据集中的3160例ICU入院病例上评估了LLaMA 3.1 8B和LLaMA 3.3 70B。与无外部知识相比,添加有临床意义的外部知识报告使8B模型的AUROC提高了8.57个百分点,70B模型提高了1.99个百分点。相对于无外部知识报告的GPT-5.2前沿模型参考(AUROC 68.86%),知识注入使LLaMA 8B的性能差距从15.66缩小到7.09 AUROC点,LLaMA 70B从5.30缩小到3.31 AUROC点。随机对照报告不能提高性能,反而常常降低性能,这表明收益取决于有临床意义的内容,而不仅仅是增加的提示长度。基于SHAP的归因进一步证实了预测过程中注入的知识被积极使用。这些发现表明,推理时的知识注入可以缩小可本地部署的开放权重模型与前沿封闭模型之间的差距,同时为资源受限的临床环境保留实用、隐私保护的工作流程。

英文摘要

Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge, especially for smaller locally deployable models. We present a lightweight knowledge-injection framework for zero-shot ICU delirium prediction that augments a deterministic natural-language summary of structured electronic health record data with an external clinical knowledge report at inference time, without fine-tuning or retrieval. We evaluate LLaMA 3.1 8B and LLaMA 3.3 70B on 3,160 ICU admissions from the MIMIC IV dataset. Adding a clinically meaningful external knowledge report improves AUROC by 8.57 percentage points for the 8B model and 1.99 percentage points for the 70B model compared to no external knowledge. Relative to a GPT-5.2 frontier-model reference without external knowledge report (AUROC 68.86%), knowledge injection reduces the performance gap from 15.66 to 7.09 AUROC points for LLaMA 8B and from 5.30 to 3.31 AUROC points for LLaMA 70B. Random control reports do not improve performance and often degrade it, indicating that gains depend on clinically meaningful content rather than added prompt length alone. SHAP-based attribution further confirms that the injected knowledge is actively used during prediction. These findings suggest that inference-time knowledge injection can narrow the gap between locally deployable open-weight models and frontier closed models while preserving a practical, privacy-preserving workflow for resource-constrained clinical settings.

↑