AI 中文总结
本研究通过训练Aether基础模型家族,揭示了生理波形基础模型的计算最优扩展规律,证明同时扩大模型规模和预训练时长可提升下游临床预测性能,并减少标注需求。
AI 中文摘要
我们研究了生理波形基础模型(FMs)的扩展规律和计算最优训练。我们训练了Aether,一个包含超过一百个基础模型的家族,参数规模从2000万到21亿不等,使用了多达3630万小时的生理波形数据。我们从MIMIC-III构建了八个临床预测任务,并通过线性探测评估这些基础模型。720M参数的基础模型在所有八个任务上均优于所有现有的基线基础模型。一个关于模型规模、预训练小时数和标注患者数量的扩展规律能够有效预测下游排序误差,即$1-\mathrm{AUROC}$,在保留的资源规模上预测平均绝对误差(MAE)为0.5%,当外推到21亿参数时MAE为0.9%。我们提出三个发现:(1)计算最优训练同时扩展基础模型规模和预训练小时数。在拟合的规律下,计算FLOPs增加10.0倍会使模型规模扩大1.2倍,预训练小时数增加8.2倍。(2)较大的基础模型更高效地利用波形数据,且更多的预训练暴露增加了模型扩展的收益。从2500万参数和480万预训练小时开始,将基础模型规模加倍可将达到相同性能所需的预测小时数减少51.8%。(3)预训练和临床监督相互增强:更多的标注患者增加了预训练的回报,而更大的基础模型和更长的预训练减少了标注需求。以720M基础模型为例,将预训练从12万小时延长到3630万小时,在目标排序误差下可将预测的患者需求减少61%。这些发现为生理波形建模和下游临床预测提供了定量的训练配方和一条有前景且持久的扩展路径。
英文摘要
We investigate the scaling laws and compute-optimal training of physiological waveform foundation models (FMs). We train Aether, a family of over one hundred FMs ranging from 20M to 2.1B parameters, on up to 36.3M hours of physiological waveforms. We construct eight clinical prediction tasks from MIMIC-III and evaluate the FMs through linear probing. The 720M FM outperforms all existing baseline FMs across all eight tasks. A scaling law of model size, pretraining hours, and labeled patients predicts downstream ranking error, i.e. $1-\mathrm{AUROC}$, effectively with $0.5\%$ prediction MAE at held-out resource scales and $0.9\%$ MAE when extrapolating to 2.1B parameters. We present three findings: (1) Compute-optimal training scales both FM size and pretraining hours. Under the fitted law, a $10.0\times$ increase in compute FLOPs scales model size by $1.2\times$ and pretraining hours by $8.2\times$. (2) Larger FMs use waveform data more efficiently, and greater pretraining exposure increases the benefit of model scaling. Starting from 25M parameters and 4.8M pretraining hours, doubling FM size reduces the predicted hours needed for the same performance by $51.8\%$. (3) Pretraining and clinical supervision reinforce each other: more labeled patients increase the return to pretraining, while larger FMs and longer pretraining reduce labeling requirements. For the example of the 720M FM, extending pretraining from 120K to 36.3M hours reduces the predicted patient requirement by $61\%$ at a target ranking error. These findings provide a quantitative training recipe and a promising and durable scaling path for physiological waveform modeling and downstream clinical prediction.