发表机构
School of Computing, Mathematics and Engineering, Charles Sturt University; Department of Computer Science, Rensselaer Polytechnic Institute(查尔斯·斯特尔特大学计算、数学与工程学院; 伦斯勒理工学院计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨用大语言模型进行零样本慢性肾病筛查的可行性,提出特征引导零样本框架,通过特征选择、序列化患者记录实现。实验表明该框架能提升模型性能,在多数据集有良好泛化性,为CKD筛查提供新方法。
AI 中文摘要
慢性肾病(CKD)的早期筛查对于预防不可逆进展至关重要。然而,许多基于机器学习的筛查方法因依赖大量标记数据集、资源密集型病理测试或高维临床特征,以及对人群和分布变化的鲁棒性有限,难以在社区和资源有限的筛查环境中部署。本研究探讨在零样本设置下使用大语言模型(LLMs)进行CKD早期筛查的可行性。我们提出了一个特征引导的零样本框架,使用一组选定的具有临床意义、易于获得的社区特征来评估LLM性能,而非详尽的临床输入。通过基于机器学习的分析进行特征选择,以识别紧凑且与临床相关的变量子集。随后,使用标准化提示模板将表格患者记录序列化为文本以实现零样本推理。使用完整特征集和选定子集评估了四个LLMs(LLaMA - 3、Qwen - 3、Mistral和GPT - 4o - mini)的零样本性能。在跨越三个国家的三个异构CKD数据集上评估了泛化性。在模型和数据集之间,选定特征集在平衡准确性和概率估计方面产生了一致且具有统计学意义的改进,达到了适合筛查目的的性能水平。这些发现表明,LLMs可以使用最少的社区可获取患者特征支持具有临床意义的、无需训练的CKD筛查,在现实世界筛查环境中为传统机器学习方法提供实际补充。
英文摘要
Early screening of chronic kidney disease (CKD) is essential for preventing irreversible progression; however, many machine learning (ML)-based screening methods remain difficult to deploy in community and resource-limited screening settings due to their reliance on large labeled datasets, resource-intensive pathology tests, or high-dimensional clinical features, and limited robustness to population and distributional shifts. This study examines the feasibility of using large language models (LLMs) for early-stage CKD screening in a zero-shot setting, without dataset-specific training. We propose a feature-guided zero-shot framework that evaluates LLM performance using a selected set of clinically meaningful, readily available community-based features, rather than exhaustive clinical inputs. Feature selection was guided by ML-based analysis to identify a compact, clinically relevant subset of variables. Tabular patient records were subsequently serialized into text using standardized prompt templates to enable zero-shot inference. The zero-shot performance of four LLMs (LLaMA-3, Qwen-3, Mistral, and GPT-4o-mini) was evaluated using both the full feature set and the selected subset. Generalizability was assessed across three heterogeneous CKD datasets spanning three countries. Across models and datasets, the selected feature set yielded consistent and statistically significant improvements in balanced accuracy and probability estimates, achieving performance levels suitable for screening purposes. These findings suggest that LLMs can support clinically meaningful, training-free CKD screening using minimal community-accessible patient features, offering a practical complement to conventional ML methods in real-world screening contexts.
CommentsAuthor-prepared preprint. The Version of Record was published in AIME 2026, Springer, and is available via the DOI
Journal refProceedings of the Artificial Intelligence in Medicine (AIME 2026)
DOI:10.1007/978-3-032-30710-1_5