发表机构
Uploading Inc; Human-centric Artificial Intelligence Centre, University of Technology Sydney; Stanford Institute for Human-Centered Artificial Intelligence (HAI), Stanford University(Uploading Inc; 悉尼科技大学以人为本人工智能中心; 斯坦福大学以人为本人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文综述脑到语言解码领域,涵盖任务、信号、方法、评估及实际应用,并提出从命令到双向认知交换的五级发展轨迹。
AI 中文摘要
脑到语言解码将语言产生、内部言语和感知相关的神经活动转化为语言或表达性输出。它为言语丧失后的沟通恢复提供了一条途径,也为研究大脑如何表征语言提供了一种手段。神经记录和表征学习的进展已将该领域从受限的识别和声学重建扩展到文本生成、流式个性化语音和面部动画。本综述综合了侵入性和非侵入性测量方面的这些发展,基于无下限年份的搜索和截至2026年9月的来源更新。我们将发音、内部和感知任务与它们所涉及的神经群体、解码器可用的表征以及这些表征所能支持的输出联系起来。我们考察了模型开发、公共资源和评估的演变,并在其报告的协议内比较了已发表的性能和通信成本。综述指出了互补的进展路径:语音、声学和语义目标保留了信息的不同方面;共享表征支持跨记录条件和任务的复用;在线通信越来越依赖于校准、反馈和用户控制以及解码准确性。共享基准支持算法比较,而纵向研究揭示了持续使用的需求。我们讨论了这些发展及其剩余局限性,然后概述了一个从命令和语言到意义、场景和双向认知交换的前瞻性五级轨迹。
英文摘要
Brain-to-language decoding translates neural activity associated with language production, internal speech and perception into linguistic or expressive outputs. It offers a route to restoring communication after speech loss and a means of studying how the brain represents language. Advances in neural recording and representation learning have expanded the field from constrained recognition and acoustic reconstruction to text generation, streaming personalised speech and facial animation. This survey synthesises these developments across invasive and non-invasive measurements, drawing on a search without a lower year limit and source-led updates through September 2026. We connect Articulated, Inner and Perceived tasks to the neural populations they engage, the representations available to decoders and the outputs those representations can support. We examine model development, public resources and the evolution of evaluation, and compare published performance and communication costs within their reported protocols. The synthesis identifies complementary routes to progress: phonetic, acoustic and semantic targets preserve different aspects of a message; shared representations support reuse across recording conditions and tasks; and online communication increasingly depends on calibration, feedback and user control alongside decoding accuracy. Shared benchmarks enable algorithmic comparisons, while longitudinal studies reveal the demands of sustained use. We discuss these developments and their remaining limitations, then outline a prospective five-level trajectory from commands and language to meaning, scenarios and bidirectional cognitive exchange