发表机构
UC Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究从fMRI信号中解码连续语言的挑战,一方面改进Huth编码管道,另一方面引入fMRIFlamingo,发现高容量语言模型本身未提升fMRI解码能力,且无严格评估会掩盖失败,强调盲控评估的重要性。
AI 中文摘要
从功能磁共振成像(fMRI)信号中解码连续语言仍然是非侵入性脑机接口研究中的核心挑战。我们进行了两项补充研究。首先,我们通过扩大体素选择(从10K到15K)、用GPT-2 medium替代GPT-1作为束搜索提议模型以及GPU加速的自助训练来改进Huth等人的岭回归编码管道,在对UTS03受试者的三个保留叙述中实现了平均METEOR = 0.149和BLEU-1 = 0.200,相对于我们的复制基线,METEOR相对增益为11%。其次,我们引入了fMRIFlamingo,它通过学习的脑分词器和感知器重采样器将血氧水平依赖(BOLD)活动映射到具有可训练门控交叉注意力层的冻结Llama-3.2-1B上。尽管在1/百排名任务中达到了42.86%的Top-1准确率,远高于随机水平,但对fMRI输入归零的盲控消融产生了几乎相同的分数,这表明明显的解码成功主要由冻结的语言先验驱动,而非神经输入。这些结果表明,高容量语言模型本身并不能提高fMRI解码能力,并且在没有严格盲控评估的情况下可能会掩盖失败。
英文摘要
Decoding continuous language from fMRI signals remains a core challenge in non-invasive brain-computer interface research. We present two complementary investigations. First, we improve the Huth et al. ridge regression encoding pipeline through expanded voxel selection (10K->15K), substitution of GPT-2 medium for GPT-1 as the beam-search proposal model, and GPU-accelerated bootstrap training, achieving mean METEOR = 0.149 and BLEU-1 = 0.200 across three held-out narratives for subject UTS03 -- an 11% relative METEOR gain over our replication baseline. Second, we introduce fMRIFlamingo, which maps BOLD activity to a frozen Llama-3.2-1B with trainable gated cross-attention layers via a learned brain tokenizer and a Perceiver Resampler. Despite achieving 42.86% Top-1 accuracy on a 1-in-100 ranking task, well above chance, a blind control ablation with zeroed fMRI inputs yields near-identical scores, revealing that apparent decoding success is driven primarily by the frozen language prior rather than by neural input. These results demonstrate that high-capacity language models do not inherently improve fMRI decoding and can actively obscure failures without rigorous blind-control evaluation.
Comments7 pages, 2 tables, 2 figures. Preprint. NLP 244: Advanced Machine Learning for Natural Language Processing report, UC Santa Cruz