用于皮层下脑机文本的音素级与字符级目标及选择性状态空间模型
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
浏览论文内容
中文总结 AI 辅助
该研究在Brain-to-Text '25基准上对比GRU与混合Mamba解码器、音素级与字符级目标的脑机文本系统,发现循环基线性能最优,混合Mamba具竞争力但未超越基线。
中文摘要 AI 辅助
当前最先进的皮层下脑机文本系统将神经序列音素解码器与外部语言模型相结合,仍有两个设计维度未得到充分探索:选择性状态空间模型(Mamba)是否优于循环解码器,以及输出目标(音素级与字符级)如何与该选择相互作用。在公开的Brain-to-Text '25基准上,我们采用可复现的训练协议,研究了受控2×2网格实验(GRU与混合Mamba解码器;音素级与字符级目标),以连接时序分类(CTC)为训练目标。循环基线仍表现最佳:最优音素级GRU达到12.62%的音素错误率(PER)和21.19%的词错误率(WER),最优文本级GRU经语言模型重评分后达到13.39%的字符错误率(CER)和26.28%的WER。混合Mamba具有竞争力但未超越基线。消融实验分离了架构贡献,误差分析显示存在依赖表示的失败情况:类发音的音素混淆与词汇及词边界错误。
英文摘要
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text '25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.
发表机构
- Universitat Oberta de Catalunya (UOC)(加泰罗尼亚开放大学)
- University of Granada(格拉纳达大学)
- Research Centre for Information and Communication Technologies (CITIC-UGR)(信息与通信技术研究中心(CITIC-UGR))
机构由 AI 辅助整理,请以论文原文为准。