发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可穿戴静默语音接口词汇量受限的问题,提出首个基于声学传感眼镜的大规模开放词汇三模态数据集SoniSpeech,构建了该任务的首个基准并取得26.3%的词错误率。
AI 中文摘要
可穿戴静默语音接口(SSI)受限于小的封闭词汇表,实现更大词汇量的方法需要使用面部电极等侵入性硬件。我们提出SoniSpeech,这是首个基于声学传感眼镜的可穿戴SSI大规模开放词汇三模态数据集,包含18000条话语共34小时数据,具备超声回波轮廓、发声音频、正面视频三种同步模态,涵盖发声与静默模式。该语料库源自SODA对话数据集,提供包含5356个独特单词、覆盖全部音素的当代英语口语。基于连接时序分类(CTC)的ResNet-34基线模型在开放词汇静默语音识别任务中达到26.3%的词错误率(WER),这是该任务的首个基准,数据集可通过指定URL获取。
英文摘要
Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear. It contains 34 hours across 18,000 utterances with three synchronized modalities: ultrasound echo profiles, voiced audio, and frontal video, in both voiced and silent modes. The corpus draws from the SODA dialogue dataset, providing contemporary conversational English with 5,356 unique words and full phoneme coverage. A CTC-based ResNet-34 baseline achieves 26.3% word error rate (WER) on open-vocabulary silent speech recognition, the first benchmark for this task. Dataset is available at https://doi.org/10.7298/xjjr-9m85