Brain2Speech-Net:无需文本解码的可理解、实时脑机语音合成
Brain2Speech-Net: Fast and Intelligible Brain-to-Speech Synthesis Without Text Decoding
浏览论文内容
中文总结 AI 辅助
该研究针对瘫痪患者言语丧失问题,提出Brain2Speech-Net单阶段框架,通过可微分音素瓶颈与深度HMM对齐器实现无需文本解码的实时可理解语音合成,性能优于级联系统。
中文摘要 AI 辅助
言语丧失会限制瘫痪患者的沟通能力,通过直接从神经活动合成语音来恢复言语极具挑战性:皮层内数据稀缺且缺乏对齐目标,因此大多数系统依赖级联的神经-文本-语音流水线,这会增加延迟并传播误差。我们提出了Brain2Speech-Net,它是首批在有限数据下仍保持可理解性且移除中间文本解码的单阶段框架之一。一个可微分音素瓶颈在无需显式文本解码的情况下保留语言结构;随后,一个轻量型深度隐马尔可夫模型(deep-HMM)对齐器将该瓶颈映射到文本转语音(TTS)潜在空间中的上下文音素表示,它无需帧级监督即可学习神经记录与音素片段之间的单调对齐,继承了强大的声学先验以实现数据高效训练。在一个皮层内数据集上,Brain2Speech-Net在客观测试和听力测试中均实现了强可理解性,同时运行速度快于实时。与延迟高的级联系统以及缺乏可理解性的直接语音单元模型不同,它同时提供可理解且实时的语音。
英文摘要
The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.
发表机构
- Johns Hopkins University(约翰霍普金斯大学)
- Academia Sinica(中央研究院)
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。