发表机构
UCLA Samueli School of Engineering; University of Pennsylvania(加州大学洛杉矶分校萨穆埃利工程学院; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PHASE提出生理引导的分层基础模型,将可测量的生理特征作为显式学习目标,在Omni-iEEG五项临床任务上以高达31%的优势超越现有基础模型,并刷新最先进水平。
AI 中文摘要
临床医生和神经科学家长期以来通过直接可测量的生理特征来分析颅内脑电图(iEEG),这些特征承载了下游任务所依赖的大部分信息。现有的iEEG基础模型通过重建或预测其输入来学习,这使得对这些特征的保留是隐性的。它们也主要是在认知解码和一个狭窄的临床任务(即癫痫发作检测)上进行评估。在诸如Omni-iEEG这样广泛且临床相关的基准上,当冻结使用时,它们仍低于特定任务模型。我们提出了PHASE,一种生理引导的基础模型,它将这些特征作为显式的学习目标,并在时间阶段(PHASE-T)中于每个通道内将它们与掩码潜变量预测配对,以及在跨同步通道的时空阶段(PHASE-ST)中进行配对。PHASE在来自九个临床站点、222名参与者的异质性记录上进行了预训练。在所有五项Omni-iEEG临床任务上,冻结的PHASE-T以高达31%的优势优于每个被评估的基础模型,而微调后的PHASE-T超越了特定任务模型,树立了新的最先进水平。PHASE-T受益于生理监督,在匹配的消融实验中,于每项任务上都优于仅使用潜变量预测或辅助波形重建训练的变体。PHASE-T能泛化到未见过的机构,在少量或无本地标签的情况下优于所比较的模型。PHASE-ST在癫痫发作起始区识别上进一步优于PHASE-T,并且当冻结时,在BrainTreebank上对声音音量和音高的解码优于已发表的模型。超越任务性能,PHASE学习封装了临床医生所识别的生理特征,从癫痫发作起始及其传播到解剖区域身份,尽管其预训练不包含任何发作期记录或解剖标签。
英文摘要
Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduce PHASE, a physiology-guided foundation model that makes these characteristics explicit learning targets, pairing them with masked latent prediction in a temporal stage (PHASE-T) within each channel and a spatiotemporal stage (PHASE-ST) across synchronized channels. PHASE is pretrained on heterogeneous recordings from 222 participants at nine clinical sites. On all five Omni-iEEG clinical tasks, frozen PHASE-T outperforms every evaluated foundation model by up to 31\%, and fine-tuned PHASE-T surpasses the task-specific models, setting a new state of the art. PHASE-T benefits from physiological supervision, outperforming variants trained with latent prediction alone or auxiliary waveform reconstruction on every task in matched ablations. PHASE-T generalizes to unseen institutions, outperforming the compared models with few or no local labels. PHASE-ST further improves seizure-onset-zone identification over PHASE-T and, when frozen, decodes sound volume and pitch on BrainTreebank better than published models. Beyond task performance, PHASE learns to encapsulate the physiological characteristics clinicians recognize, from seizure onset and its propagation to anatomical region identity, even though its pretraining contains no ictal recordings or anatomical labels.
CommentsAuthor Metadata Correction