AI 中文总结
本文从系统视角综述语音脑机接口,强调其作为自适应临床系统的本质,并指出下一代评估需兼顾离线准确性与跨会话稳健性、延迟、可用性等多维指标,以推动可靠可控的通信神经假体发展。
AI 中文摘要
语音脑机接口(BCIs)旨在通过将与语音、语言或交际意图相关的神经活动转化为外部输出(如文本、合成语音或虚拟形象控制)来恢复沟通能力。近年来,皮层内和皮层脑电图记录、深度序列模型以及语言模型辅助解码的进展推动了该领域的快速发展,包括高性能的尝试性语音解码和日益自然的语音合成。然而,这些成就也揭示出语音脑机接口并非简单的神经到文本解码器。它们是自适应的临床系统,其中神经表征、记录硬件、解码架构、语言先验、反馈和用户学习随时间相互作用。在此,我们从系统级视角综合审视语音脑机接口研究。首先,我们考察语音和语言的神经基础,强调其分层、分布式、时间结构化且非平稳的组织特性。随后,我们审视记录与解码选择、闭环自适应、评估、临床转化及伦理问题。在这些领域中,我们强调信号分辨率与侵入性、低层运动目标与高层语义目标、解码器准确性与用户自主性、以及语言模型流畅性与忠实神经证据之间反复出现的权衡。我们认为,下一代语音脑机接口不仅应通过离线准确性进行评估,还应考虑跨会话的稳健性、校准负担、延迟、不确定性、可用性以及针对非预期解码的防护措施。通过将语音脑机接口重新定义为自适应、以用户为中心的系统,我们勾勒出跨越语音神经科学、神经工程、机器学习、临床实践和神经伦理学的跨学科优先事项,以推动从概念验证解码迈向可靠、富有表现力且可控的通信神经假体。
英文摘要
Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.
CommentsReview article, 28 pages, 4 figures, 2 boxes, 2 tables