发表机构
Tether Evo; Turner Institute for Brain and Mental Health, Monash University; The University of Rome Tor Vergata(Tether Evo; 莫纳什大学特纳大脑与心理健康研究所; 罗马第二大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BRAID-fMRI框架,利用共享CLIP监督的ROI-wise Transformer,在八个视觉fMRI数据集上实现跨参与者与数据集的统一解码,显著提升检索准确率并揭示腹侧视觉皮层的关键作用。
AI 中文摘要
从fMRI进行视觉解码通常受限于参与者和实验的孤立性,这掩盖了异质性神经测量能否被组织在共同的计算几何中的问题。在此,我们引入BRAID-fMRI(脑表征跨个体与数据集对齐),一种共享的CLIP监督解码框架。BRAID-fMRI使用单一的全脑区域(ROI-wise)Transformer,可选择性地对八个视觉fMRI数据集进行参与者条件化,这些数据集包含93个数据集特定的参与者条目、430,007个单次试验响应和162,839个独特刺激。区域脑活动通过多正样本对比目标与512维CLIP ViT-B/32表征对齐,该目标将跨参与者和数据集的重复刺激视为正样本。BRAID-fMRI支持跨七个评估数据集的检索。在八个匹配的参与者条目上,它实现了35.0 ± 11.1%的Top-10准确率,超过了两个评估基线——MindEye风格的池化CLIP解码器(27.1 ± 5.1%)和岭回归(21.1 ± 9.1%)——的观测平均准确率,并在八个条目中的七个上达到最高观测准确率。在分别训练的参与者无关模型中,扩大源池相对于初始源训练条件,将目标数据集留出准确率提高了最多92.1%。学习到的空间保留了分级语义结构,而消融和显著性分析突出了腹侧和早期视觉皮层以及类别特定的运动和注意系统。这些结果支持将可扩展的跨数据集解码纳入共同的CLIP对齐空间,模型敏感性集中在腹侧和早期视觉输入上。
英文摘要
Visual decoding from fMRI is typically siloed by participant and experiment, obscuring whether heterogeneous neural measurements can be organized within a common computational geometry. Here we introduce BRAID-fMRI (Brain Representation Alignment across Individuals and Datasets), a shared CLIP-supervised decoding framework. BRAID-fMRI uses a single ROI-wise Transformer with optional participant conditioning across eight visual-fMRI datasets comprising 93 dataset-specific participant entries, 430,007 single-trial responses and 162,839 unique stimuli. Regional brain activity is aligned with 512-dimensional CLIP ViT-B/32 representations using a multi-positive contrastive objective that treats repeated stimuli across participants and datasets as positives. BRAID-fMRI supports retrieval across seven evaluation datasets. On eight matched participant entries, it achieves 35.0 +/- 11.1% Top-10 accuracy, exceeding the observed mean accuracy of the two evaluated baselines - the MindEye-style pooled-CLIP decoder (27.1 +/- 5.1%) and ridge regression (21.1 +/- 9.1%) - and attaining the highest observed accuracy for seven of eight entries. In separately trained participant-agnostic models, expanding the source pool increased target-dataset holdout accuracy by up to 92.1% relative to the initial source-training condition. The learned space preserves graded semantic structure, while ablations and saliency highlight ventral and early visual cortex and category-specific motion and attentional systems. These results support scalable cross-dataset decoding into a common CLIP-aligned space, with model sensitivity concentrated in ventral and early visual inputs.
Comments21 pages, 5 figures, 1 table, 3 supplementary figures, 2 supplementary tables