现象图JEPA:用于非接触式心肺传感的标签高效表示学习
Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing
浏览论文内容
中文总结 AI 辅助
本文提出现象图JEPA,一种无需负样本的联合嵌入预测架构,用于非接触式心肺传感的标签高效表示学习,在OMuSense-23上显著提升标签效率,并区分了预测表示的收益与生理先验的作用。
中文摘要 AI 辅助
毫米波雷达和RGB-D相机可以连续且非接触地记录心脏和呼吸波形,但由于每个标签都需要一次有监督的采集过程,带标签的记录仍然稀缺。自监督预训练可以利用未标记的信号,然而对比方法依赖于信号变换和负样本对,其有效性对于心肺数据而言是不确定的,因为时间扭曲会改变呼吸频率,而相隔较远的窗口可能共享相同的生理状态。我们提出了现象图JEPA,一种联合嵌入预测架构,其基础配置无需负样本对或合成增强即可从四个处理过的一维数据流中学习。每个数据流由一个时间卷积分支和一个带限频谱分支编码。在预训练期间,模型沿着类型化边预测被掩码的目标嵌入,这些边连接被分配到同一生理现象的数据流,并在状态片段内随时间向前推进。我们将这种生理类型化视为一个可检验的假设,并将其与错误边和全对预测图进行比较。在OMuSense-23数据集中,预训练将标签效率面积相对于匹配的有监督训练提高了3.91个百分点(95%置信区间为2.08至5.80,Holm校正后p=0.006),在第二种配置下,对相同测试参与者评估时提高了3.74个百分点。然而,错误边和全对控制并未证明生理类型化的益处。可选的基于Takens思想的延迟坐标改善了与学习历史相比的验证比较,而两个仅使用腕部传感器的WESAD协议并未证明预训练优势。因此,该研究将预测表示的可测量益处与用于组织其训练的生理先验区分开来。
英文摘要
Millimeter-wave (mmWave) radar and RGB-D cameras can record cardiac and respiratory waveforms continuously and without contact, but labeled recordings remain scarce because every label requires a supervised acquisition session. Self-supervised pretraining can exploit the unlabeled signals, yet contrastive methods depend on signal transformations and negative pairs whose validity is uncertain for cardiorespiratory data, where time warping changes breathing rate and distant windows can share the same physiological state. We present Phenomenon-Graph JEPA, a joint-embedding predictive architecture that learns from four processed one-dimensional streams without negative pairs or synthetic augmentation in its base configuration. Each stream is encoded by a temporal convolutional branch and a band-limited spectral branch. During pretraining, the model predicts stopped target embeddings along typed edges, which connect streams assigned to the same physiological phenomenon, and forward in time within a state episode. We treat this physiological typing as a testable hypothesis and compare it with wrong-edge and all-pairs prediction graphs. In the OMuSense-23 dataset, pretraining improves label-efficiency area over matched supervised training by 3.91 percentage points (95% interval 2.08 to 5.80, Holm-adjusted p = 0.006), and by 3.74 points under a second configuration evaluated on the same test participants. However, the wrong-edge and all-pairs controls do not establish a benefit from physiological typing. Optional Takens-inspired delay coordinates improve a validation comparison with learned history, whereas two wrist-only WESAD protocols do not establish a pretraining advantage. The study therefore separates the measured benefit of predictive representations from the physiological prior used to organize their training.