发表机构
University of Edinburgh; Health Data Research UK; University of Toronto; Western University(爱丁堡大学; 英国健康数据研究中心; 多伦多大学; 西安大略大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ViSTA通过紧凑适配器将数值测量融入视觉-语言模型,实现临床时间序列预测与时间问答,在MIMIC-IV上以极少参数达到高准确率。
AI 中文摘要
临床预测模型根据患者测量数据估计风险,而大语言模型支持医学文本理解和问答。然而,它们的语言能力并不能确保从结构化、高维的临床时间序列中进行准确预测。提升这一能力将把风险估计与关于患者病情演变的灵活问题联系起来。我们引入了ViSTA,一个紧凑的适配器,它将不规则的数值测量融入预训练的视觉-语言模型的图表表示中。它学习对视觉令牌的修正,同时保持所有预训练参数不变。在MIMIC-IV上,ViSTA在2-9亿参数的模型中,针对急性肾损伤和死亡率预测的所有四个指标上,其平均得分均高于对比的适配方法。凭借51.6万可训练参数,20亿参数的模型在急性肾损伤上达到了0.7376的ROC曲线下面积,而GPT-5.6 Sol在文本输入和高推理努力下为0.7380。针对时间问答的训练在40亿参数下达到了69.27%的准确率,其可训练参数比使用图表或数值文本的低秩适配减少了90%以上,准确率差距为2.82-4.88个百分点。ViSTA将预训练语言模型扩展到数值预测和时间问题。
英文摘要
Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from structured, high-dimensional clinical time series. Improving this ability would connect risk estimation with flexible questions about a patient's evolving condition. We introduce ViSTA, a compact adapter that incorporates irregular numerical measurements into a pretrained vision-language model's chart representations. It learns corrections to visual tokens while leaving all pretrained parameters unchanged. On MIMIC-IV, ViSTA has the highest mean scores among the compared adaptations on all four metrics for acute kidney injury and mortality prediction across models with 2-9 billion parameters. With 0.516 million trainable parameters, the 2-billion-parameter model reaches an area under the ROC curve of 0.7376 for acute kidney injury, compared with GPT-5.6 Sol's 0.7380 with text input and high reasoning effort. Training for temporal question answering yields 69.27% accuracy at 4 billion parameters with over 90% fewer trainable parameters than low-rank adaptation using charts or numerical text, at a 2.82-4.88 percentage-point accuracy gap. ViSTA extends pretrained language models to numerical prediction and temporal questions.