发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MEG-Mamba,一种基于Mamba架构的脑磁图生成式基础模型,通过预测离散化信号的下一个标记实现高效预训练,在更短时间和更长上下文下超越Transformer模型,并支持轻量级刺激条件化生成任务诱发反应,表明Mamba是扩展神经基础模型的可行主干。
AI 中文摘要
脑磁图(MEG)是一种成像技术,能够以非侵入性、毫秒级分辨率观察人类大脑活动。MEG数据的日益普及为利用人工智能领域的最新进展(即自监督基础模型)提供了机会。现有的MEG(和脑电图)基础模型基于Transformer架构构建。在此,我们引入MEG-Mamba:一种基于Mamba架构的神经活动(源重建、分区MEG)生成式基础模型。MEG-Mamba经过训练,能够根据大脑区域和记录会话的条件,预测离散化MEG信号的下一个标记。MEG-Mamba在生成保真度上超越了基于Transformer的替代模型(MEG-GPT),同时预训练时间更短(22 GPU小时对比400 GPU小时),并建模更长的上下文(4秒对比0.32秒)。我们通过检查MEG-Mamba生成数据的时空频谱特征以及解释其学习到的嵌入来评估该模型。此外,我们证明轻量级刺激条件化(LoRA)可用于生成预训练中未包含的真实任务诱发反应。我们的结果表明,Mamba是扩展神经基础模型的有前景的主干网络。
英文摘要
Magnetoencephalography (MEG) is an imaging technique that offers a non-invasive, millisecond-resolution view of human brain activity. The increasing availability of MEG data presents an opportunity to take advantage of a recent advance in artificial intelligence, namely self-supervised foundation models. Existing foundation models for MEG (and electroencephalography) have been built on a transformer architecture. Here, we introduce MEG-Mamba: a generative foundation model for neural activity (source reconstructed, parcellated MEG) built on a Mamba architecture. MEG-Mamba is trained to predict the next token of a discretised MEG signal, conditioned on a brain region and recording session. MEG-Mamba surpasses the generative fidelity of a transformer-based alternative (MEG-GPT) while pre-training in less time (22 vs 400 GPU-hours) and modelling a longer context (4 vs 0.32 s). We evaluate MEG-Mamba by examining the spatio-spectral characteristics of the data it generates and by interpreting its learned embeddings. Furthermore, we demonstrate that lightweight stimulus conditioning (LoRA) can be used to generate realistic task-evoked responses that are not included in pre-training. Our results suggest that Mamba is a promising backbone for scaling neural foundation models.
Comments15 pages, 5 figures