arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SENSE:通过空间图编码从脑动态中实现语义神经语音合成

SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding

Jisoo Park, Seonghak Lee, Hyojin Park, Junseok Kwon

arXiv 2609.37601首次发表:更新:

发表机构

Chung-Ang University; University of Birmingham(中央大学; 伯明翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对脑电图到语音合成忽略语义的问题,提出SENSE方法,利用图编码和语义条件对齐,在N400数据集上优于现有方法,且小样本下表现优异。

AI 中文摘要

从非侵入性脑信号中重建语音,为恢复认知完整但无法说话的个体的沟通提供了一条有前景的途径。现有的脑电图到语音方法将此任务表述为声学重建,优化波形保真度,而忽略了生成的语音是否保留高层语义内容。在本工作中,我们重新审视了这一表述,并认为脑电图信号不仅携带声学信息,还携带语义信息。我们识别出先前方法的两个关键局限性:(1)忽略脑电图电极之间的空间关系,(2)未能利用N400范式的语义结构,其中一致和不一致的试验反映了不同的语义处理。我们提出SENSE(语义脑电图神经语音合成),它结合了基于图的脑电图编码器(基于电极几何)与脑电图语义条件(ESC),仅使用一致试验将脑电图对齐到预训练的语义空间。在N400数据集上,SENSE在声学和语义指标上均持续优于先前方法,且模型内部的通道归因表明其对听觉、感觉运动和中央-顶叶区域的分布式依赖,这与已知的语音感知神经科学一致。在未见受试者设置中,仅用两个受试者训练的SENSE在词错误率上已经超过了用全部十八个受试者训练的最强基线,并且仅用八个受试者就在声学指标上与之匹配。

英文摘要

Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as acoustic reconstruction, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose SENSE(Semantic-EEG Neural Speech SynthEsis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.

CommentsAccepted at NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑