Myovox:从面部肌肉读取语音
Myovox: Reading Speech from the Muscles of the Face
浏览论文内容
中文总结 AI 辅助
Myovox利用面部31通道sEMG信号解码开放词汇英语,通过恢复基线、双向Conformer蒸馏和LLM重排序三步,将词错误率从51.17%降至18.53%,并揭示声学音素错误率是性能瓶颈。
中文摘要 AI 辅助
Myovox,取自myo(肌肉)和vox(声音),从发声言语期间面部肌肉记录的31通道表面肌电信号(sEMG)解码开放词汇英语文本。它将已发表的单受试者emg2speech General Corpus的词错误率从51.17%降至18.53%,通过三个可分离的步骤实现,每一步均单独测量。首先,我恢复了公开版本中缺失的开放词汇解码设置,并达到了忠实的40.63% WER / 39.02% PER基线,其音素错误率与已发表结果相差在0.8个百分点以内,因此声学模型被忠实地复现。其次,我将因果编码器替换为双向Conformer,通过四项交叉模态蒸馏针对并行音频的WavLM-Large第9层特征进行训练,仅凭肌电信号即达到26.14% WER / 22.34% PER。第三,我将两个声学模型集成,合并其多尺度n-best列表,并使用QLoRA微调的7B语言模型进行重排序,达到18.53% WER,这是该语料库上报告的最佳结果,尽管并非其他语料库上sEMG到文本的最佳报告结果(第2节)。随后,我报告了限制整个方法的负面结果:重排序在18.5%处已耗尽,因为约束因素是肌电声学音素错误率(约20.9%),而非语言模型。正确的单词根本不存在于声学后验中,因此任何重排序器都无法达到9.30%的n-best oracle。所有测试数字均基于400句保留测试集,采用作者官方的8,500 / 760 / 400顺序划分;每个超参数仅在验证集上调整一次,并应用于测试集一次。
英文摘要
Myovox, from myo (muscle) and vox (voice), decodes open-vocabulary English text from 31-channel surface electromyography (sEMG) recorded from the muscles of the face during vocalized speech. It takes the single-subject emg2speech General Corpus from a published 51.17% word error rate to 18.53%, in three separable moves, each measured in isolation. First, I recover the open-vocabulary decode settings missing from the public release and reach a faithful 40.63% WER / 39.02% PER baseline whose phone error rate matches the published one to within 0.8 points, so the acoustic model is reproduced faithfully. Second, I replace the causal encoder with a bidirectional Conformer trained by a four-term cross-modal distillation against the parallel audio's WavLM-Large layer-9 features, reaching 26.14% WER / 22.34% PER from the electromyography alone. Third, I ensemble two acoustic models, union their multi-scale n-best lists, and rerank with a QLoRA-fine-tuned 7B language model, reaching 18.53% WER, the best result reported on this corpus, though not the best reported for sEMG-to-text on other corpora (Section 2). I then report the negative result that bounds the whole approach: reranking is exhausted at 18.5% because the binding constraint is the electromyographic acoustic phone error rate (~20.9%), not the language model. The correct words are simply absent from the acoustic posteriors, so no reranker can reach the 9.30% n-best oracle. All test numbers are on the 400-sentence held-out test set under the authors' official 8,500 / 760 / 400 sequential split; every hyperparameter is tuned once on validation and applied once to test.