独立大型语言模型(LLM)与预定义智能体管道用于解释ICU死亡率预测:基于eICU演示数据集的可行性研究
Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset
浏览论文内容
中文总结 AI 辅助
本研究基于eICU演示数据集对比独立LLM与四步智能体管道解释ICU死亡率预测,发现智能体管道在指南依据性等方面更优,但需搭配归因核查。
中文摘要 AI 辅助
机器学习模型可准确预测ICU死亡率,但仅特征归因方法很少能提供床边使用所需的临床叙述。大型语言模型(LLM)或可弥合这一差距,多步骤智能体管道是合理扩展,因其可分离数据解释、指南核查与最终解释。本修订可行性研究保留了原有的独立模型与智能体管道的对比,同时使主要临床发现更明确。使用保留的本地eICU演示人工制品集(2353次ICU住院,死亡率8.1%),XGBoost的AUROC为0.855(95%置信区间0.796--0.906),AUPRC为0.332(95%置信区间0.217--0.494)。在分层的38例解释子集上,独立LLM产生1例存在明确结果泄露的解释,而四步智能体管道无此类情况。在与SHAP审查子集重叠的14例中,独立LLM的SHAP对齐度更高(平均Jaccard系数0.171对比0.077)、方向一致性更高(92.9%对比78.6%),而智能体管道的指南依据性更高(0.762对比0.143)、值特异性更高(0.236对比0.143),合理性略高(0.700对比0.671)。临床层面,结果表明智能体分解或可提升与安全相关的依据性及患者特异性细节,但在高风险解释场景使用前,应搭配基于归因的核查。
英文摘要
Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation. This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit. Using the retained local eICU Demo artifact set (2,353 ICU stays; 8.1\% mortality), XGBoost achieved an AUROC of 0.855 (95\% CI 0.796--0.906) and an AUPRC of 0.332 (95\% CI 0.217--0.494). On a stratified 38-case explanation subset, the standalone LLM produced 1 explanation with explicit outcome leakage, whereas the four-step agentic pipeline produced none. Among the 14 cases that overlapped with the SHAP review subset, the standalone LLM showed higher SHAP alignment (mean Jaccard 0.171 versus 0.077) and higher direction consistency (92.9\% versus 78.6\%), while the agentic pipeline showed higher guideline grounding (0.762 versus 0.143), higher value specificity (0.236 versus 0.143), and slightly higher plausibility (0.700 versus 0.671). Clinically, the results suggest that agentic decomposition may improve safety-relevant grounding and patient-specific detail, but it should be paired with attribution-based checks before use in high-stakes risk explanation.
发表机构
- Santa Clara University(圣克拉拉大学)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
- University of Pennsylvania(宾夕法尼亚大学)
- Carnegie Mellon University(卡内基梅隆大学)
- New York University(纽约大学)
- Wake Forest University(维克森林大学)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。