发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示非侵入式脑到文本解码中的时间捷径,通过独立处理窗口移除该捷径,显著提升解码性能,接近侵入式方法。
AI 中文摘要
我们发现,从非侵入性脑记录中解码单词的主要报道改进,在很大程度上无需任何脑数据即可复现。在d'Ascoli等人(2025)的有影响力的工作中,受试者感知连续语音时的脑活动时间序列被分割成从每个单词开始的固定长度窗口。随后,一个神经网络同时生成句子中所有单词的预测。相邻窗口部分重叠,隐式揭示了单词之间的间隔。由于这些间隔指示了所说单词的时长,而不同单词往往具有不同的时长——例如,“the”比“supercalifragilisticexpialidocious”短得多——神经网络可以在不依赖底层脑活动的情况下改进其单词预测。与此一致,该方法在包含无脑信息的合成信号上达到22.0%的平衡准确率,而在真实脑记录上为22.3%。为防止网络学习这一捷径,我们做了一个简单改变。不再联合编码句子中的所有窗口,而是独立处理每个窗口。结果,神经网络通过从脑记录中学习底层单词特定信息,实现了更好的性能。这使得两种现有策略比之前有效得多。对同一单词的不同神经反应进行预测聚合,以及使用预训练LLM作为语言先验,现在都大幅提升了结果。在我们的感知语音基准上,这一简单配方(SimpleB2T)在每词五次观测下实现了36.6%的词错误率,接近过去的侵入式语音解码性能,尽管条件不同。本工作的结果暴露了脑到文本解码中的一个重要捷径,并表明移除它导致了一个简单且更有效的策略。
英文摘要
We find that major reported improvements in decoding words from non-invasive brain recordings are largely reproducible without any brain data. In the influential work of d'Ascoli et al. (2025), time series of brain activity from subjects perceiving continuous speech are segmented into fixed-length windows starting at each word. A neural network then generates predictions for all of the words in a sentence together. Neighbouring windows partially overlap, implicitly revealing the interval between words. Since these intervals indicate the duration of the words spoken, and different words tend to have different durations - for example, "the" is much shorter than "supercalifragilisticexpialidocious" - the neural network can improve its predictions of words without relying on the underlying brain activity. Consistent with this, the method reaches 22.0% balanced accuracy on synthetic signals containing no brain information, compared with 22.3% on real brain recordings. To prevent the network from learning this shortcut, we make a single, simple change. Instead of jointly encoding all windows in a sentence, we process each independently. As a result, the neural network achieves better performance by learning underlying word-specific information from brain recordings. This makes two existing strategies become much more effective than before. Both aggregating predictions from distinct neural responses to the same word and using a pretrained LLM as a linguistic prior now substantially improve results. On our perceived speech benchmark, this simple recipe (SimpleB2T) achieves a word error rate of 36.6% with five observations per word, approaching past invasive speech decoding performance, albeit under different conditions. The results in this work expose an important shortcut in brain-to-text decoding and show that removing it leads to a simple and considerably more effective strategy.
Comments29 pages, 12 figures, 10 tables