arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21288cs.LG

多受试者预训练实现封闭语料库表面肌电语音解码的短校准个性化

Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

Chenqian Le, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Tianyu He, Nikasadat Emami, Adeen Flinker, Yao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过多受试者预训练和检查点初始化,在封闭语料库中实现表面肌电语音解码的短校准个性化,显著降低字符和词错误率。

中文摘要 AI 辅助

基于表面肌电(sEMG)的无声语音接口受到跨用户差异性和校准负担的限制。我们研究了一种有限数据设置,其中27名具有典型言语能力的参与者,在朗读(Aloud)和默读(Mimed)两种语音模式下,每人贡献了少于0.5小时的数据(平均21.3分钟)。在封闭的50句语料库内,我们采用留一受试者(leave-one-subject-out)评估方法,从已发布的单受试者检查点初始化,在非留出参与者上进行预训练,并在目标参与者上进行微调。该流程实现了21.7%的字符错误率(CER)和31.9%的词错误率(WER),而未经目标受试者校准的CER为49.3%,直接检查点微调的CER为68.0%。从随机初始化开始的多受试者预训练后微调达到44.9%的CER,并且在27折中的5折中未能在固定调度下收敛,这表明检查点初始化带来了显著的优化和准确性优势。宏平均CER从使用1个预训练参与者时的74.4%下降到使用26个时的21.7%。3分钟的目标受试者校准实现了20.5%的CER和31.7%的WER,与完整的约13分钟数据池(21.7% CER和31.9% WER)相比无统计学显著差异。受试者特定适配器未提供可检测的益处。将所有sEMG模型训练数据中排除5个评估句子后,CER和WER分别增加到78.6%和99.9%。这些结果支持在标准化导联、封闭语料库设置中的短校准个性化。

英文摘要

Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5 h of data (21.3 min on average) across Aloud and Mimed speech. Within a closed 50-sentence corpus, we used leave-one-subject-out evaluation, initializing from a released single-subject checkpoint, pretraining on non-held-out participants, and fine-tuning on the target participant. This pipeline achieved 21.7% character error rate (CER) and 31.9% word error rate (WER), compared with 49.3% CER without target-subject calibration and 68.0% CER for direct checkpoint fine-tuning. Multi-subject pretraining from random initialization followed by fine-tuning reached 44.9% CER and did not converge under the fixed schedule in 5 of 27 folds, indicating substantial optimization and accuracy benefits from checkpoint initialization. Macro-averaged CER declined from 74.4% with one pretraining participant to 21.7% with 26. Three minutes of target-subject calibration achieved 20.5% CER and 31.7% WER, with no statistically significant difference from the full approximately 13-min pool (21.7% CER and 31.9% WER). A subject-specific adapter provided no detectable benefit. Excluding the five evaluation sentences from all sEMG model-training data increased CER and WER to 78.6% and 99.9%. These results support short-calibration personalization in a standardized-montage, closed-corpus setting.

发表机构

  • New York University Tandon School of Engineering(纽约大学坦登工程学院)
  • New York University Grossman School of Medicine(纽约大学格罗斯曼医学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑