发表机构
Emory University; University of California, Los Angeles; Google; Stanford University; University of Southern California; University of Wisconsin–Madison(埃默里大学; 加州大学洛杉矶分校; 谷歌; 斯坦福大学; 南加州大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PPG-LM是首个PPG-语言模型家族,通过两阶段框架对齐信号与临床文本,在约7.3万小时数据上预训练,提升检索、标题事实性和临床预测性能。
AI 中文摘要
光电容积描记(PPG)被临床监护仪和消费级可穿戴设备广泛记录,提供了可扩展的连续生理信息源。这些记录为大规模生理评估提供了机会,但实现这一潜力需要模型从信号衍生的生理监督以及电子健康记录(EHRs)中捕获的更广泛临床背景中学习。这涉及将局部观测、护理事件和整个就诊的信息与相应时间尺度的PPG表示进行对齐。然而,现有的PPG基础模型主要依赖任务特定的预测头,而大语言模型的医学知识并不一定能转化为波形理解。为弥合这一差距,我们引入了PPG-LM,这是首个从信号衍生监督和EHRs中捕获的更广泛临床背景中学习生理表征的PPG-语言模型家族。为了构建临床依据的标题,我们开发了一个自动标题生成流程,从信号测量和结构化EHR记录中生成段级、事件级和就诊级描述。然后,我们通过一个两阶段框架从这些配对中学习,该框架首先通过对比学习和波形条件标题生成建立段-语言对应关系,然后通过时间感知聚合和时间语句匹配将对齐扩展到事件和就诊。在约7.3万小时PPG上预训练,PPG-LM支持基于语言的识别、跨模态检索和段标题生成。在MC-MED、MIMIC-III和VitalDB上的实验显示,与语言模型基线相比,检索和标题事实性有所提高,并且在多个临床预测任务上优于PPG和时间序列基础模型。
英文摘要
Photoplethysmography (PPG) is widely recorded by clinical monitors and consumer wearables, providing a scalable source of continuous physiological information. These recordings offer an opportunity for physiological assessment at scale, but realizing this potential requires models to learn from both signal-derived physiological supervision and broader clinical context captured in electronic health records (EHRs). This involves aligning information spanning local observations, care events, and entire visits with PPG representations at corresponding temporal scales. However, existing PPG foundation models primarily rely on task-specific prediction heads, while the medical knowledge of large language models does not necessarily translate into waveform understanding. To bridge this gap, we introduce PPG-LM, the first PPG-language model family to learn physiological representations from both signal-derived supervision and broader clinical context captured in EHRs. To construct clinically grounded captions, we develop an automatic captioning pipeline that generates segment-, event-, and visit-level descriptions from signal measurements and structured EHR records. We then learn from these pairs through a two-stage framework that first establishes segment-language correspondence through contrastive learning and waveform-conditioned captioning, then extends alignment to events and visits through time-aware aggregation and temporal statement matching. Pretrained on approximately 73k hours of PPG, PPG-LM supports language-based recognition, cross-modal retrieval, and segment captioning. Experiments on MC-MED, MIMIC-III, and VitalDB show improved retrieval and caption factuality over language-model baselines and gains over PPG and time-series foundation models on multiple clinical prediction tasks.