发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对非侵入式语音解码信噪比低的问题,提出Brain2Semantics2Text方法,通过中间语义嵌入空间重建文本,利用语义瓶颈恢复高级语义,无需单词级对齐,并在句子级别结果上优于先前方法。
AI 中文摘要
非侵入式语音解码仍受限于神经记录的低信噪比,这使得对音素或单个单词的细粒度重建变得困难。受神经科学证据的启发,即高级语义表示分布在大脑皮层区域并以较慢的时间尺度演化,我们假设语义内容可能比低级声学或词汇特征更适合作为非侵入式解码的目标。我们提出了Brain2Semantics2Text,一种通过中间语义嵌入空间重建文本的方法。我们的模型将句子级别的脑磁图(MEG)响应映射到语义流形中,然后将预测的嵌入逆转为自然语言。这种语义瓶颈能够在无需单词级对齐的情况下恢复高级语义。我们描述了该方法的核心原理、实现方式以及用于缓解学习可靠的神经到语义映射所面临挑战的策略。最后,我们与之前的非侵入式Brain2Text方法进行了比较,并展示了在句子级别结果上的改进。
英文摘要
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.
Comments12 pages, 8 figures