发表机构
Southern University of Science and Technology; Shenzhen University; University of Warwick; SLAI(南方科技大学; 深圳大学; 华威大学; SLAI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受控实验揭示fMRI基础模型的扩展规律,发现数据和模型规模应同步扩展,且数据增加比模型规模更有效,并据此在固定计算预算下选出性能最优的模型组合。
AI 中文摘要
扩展定律已指导计算机视觉和自然语言处理领域的大模型开发,但对于功能性磁共振成像(fMRI)基础模型而言,数据、模型规模和计算量之间的关系仍不明确。在此,我们利用来自200多个源数据集的预训练数据和超过10,000个GPU小时的实验,进行了一项受控的实证研究。在保持预训练框架和下游协议固定的情况下,我们变化预训练数据规模、模型规模和训练时长。下游性能通常随计算量增加而提升,但使用相似计算量的模型表现可能差异显著。额外的预训练数据在更大模型规模下带来更大收益,这表明数据和模型规模应同步扩展。在计算量匹配的情况下,增加预训练数据比增加模型规模能惠及更多任务,尽管这种模式因任务而异。随后,我们使用分布内(ID)下游性能,在两个固定的计算预算下选择预训练数据规模、模型规模和训练时长的组合。所得模型在分布外(OOD)评估之前被锁定。在比较的fMRI基础模型中,它们在使用更少预训练计算量的同时,在评估的OOD任务上取得了最高的平均性能。总体而言,我们的结果表明,仅计算量并不能表征fMRI的扩展规律:性能取决于预训练数据、模型规模和训练时长如何组合。
英文摘要
Scaling laws have guided large-model development in computer vision and natural language processing, but the relationships among data, model size, and compute remain unclear for functional magnetic resonance imaging (fMRI) foundation models. Here, we conduct a controlled empirical study using pretraining data from more than 200 source datasets and over 10,000 GPU-hours of experiments. Holding the pretraining framework and downstream protocol fixed, we vary pretraining data size, model size, and training duration. Downstream performance generally improves with compute, yet models using similar compute can perform substantially differently. Additional pretraining data bring larger gains at larger model sizes, suggesting that data and model size should be scaled together. At matched compute, increasing pretraining data benefits more tasks than increasing model size, although the pattern varies across tasks. We then use in-distribution (ID) downstream performance to select the combination of pretraining data size, model size, and training duration at two fixed compute budgets. The resulting models are locked before out-of-distribution (OOD) evaluation. They achieve the highest average performance across the evaluated OOD tasks among the compared fMRI foundation models while using less pretraining compute. Overall, our results show that compute alone does not characterize fMRI scaling: performance depends on how pretraining data, model size, and training duration are combined.
Comments28 pages, 7 figures. Code: https://github.com/derrz2/neurojepa