关键在于表达方式:信息表示在基于LLM的血糖事件预测中的作用
It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction
浏览论文内容
中文总结 AI 辅助
本研究通过OhioT1DM数据集评估基于提示的LLMs在1型糖尿病血糖事件预测中的表现,发现信息表示方式显著影响低血糖预测性能,而传统监督模型在高血糖预测上更优。
中文摘要 AI 辅助
大型语言模型(LLMs)正越来越多地被研究用于生理时间序列预测,然而其有效性可能不仅取决于模型本身,还取决于生理信息在推理时如何被表示和呈现。本研究探讨了基于提示的通用LLMs在1型糖尿病患者餐后高血糖和低血糖预测中的应用。利用OhioT1DM数据集,我们在30、60和90分钟的预测时域内,对多个开放权重LLMs进行了零样本和少样本推理评估。分析同时变化了可用生理信息的文本表示方式以及暴露给模型的信息量,范围从仅血糖观测值到衍生描述符以及与胰岛素、膳食、碳水化合物和体力活动相关的额外上下文变量。性能与传统的患者特定监督模型以及Gluco-LLM(一种专门针对血糖时间序列预测而调整的基于语言模型的架构)进行了比较。结果显示明显的任务依赖性行为。传统的监督模型在高血糖预测中达到最强性能,而观察到的基于提示的最佳LLM配置在所有研究的时域内均改善了低血糖预测性能。基于提示的推理的有效性也受到生理信息表示方式的强烈影响,而提供额外的上下文信息并未带来系统性的改进。总体而言,这些发现强调了生理信息表示是基于提示的LLM方法在血糖事件预测中的一个核心设计因素。
英文摘要
Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only on the model itself, but also on how physiological information is represented and presented at inference time. This study investigates prompt-based general-purpose LLMs for postprandial hyperglycemia and hypoglycemia prediction in individuals with type 1 diabetes. Using the OhioT1DM dataset, we evaluate multiple open-weight LLMs under zero-shot and few-shot inference across prediction horizons of 30, 60, and 90 minutes. The analysis varies both the textual representation of the available physiological information and the amount of information exposed to the model, ranging from glucose observations alone to derived descriptors and additional contextual variables related to insulin, meals, carbohydrates, and physical activity. Performance is compared with conventional patient-specific supervised models and with Gluco-LLM, a language-model-based architecture explicitly adapted to glucose time-series forecasting. Results show a marked task-dependent behavior. Conventional supervised models achieve the strongest performance for hyperglycemia prediction, whereas the best observed prompt-based LLM configurations improve performance for hypoglycemia across all investigated horizons. The effectiveness of prompt-based inference is also strongly influenced by how physiological information is represented, while providing additional contextual information does not lead to a systematic improvement. Overall, these findings highlight physiological information representation as a central design factor in prompt-based LLM approaches to glycemic-event prediction.
发表机构
- University of Salerno(萨勒诺大学)
- University of Naples Federico II(那不勒斯费德里科二世大学)
机构由 AI 辅助整理,请以论文原文为准。