arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向儿童语音的边缘音素识别:基于年龄感知训练

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

Matthew Arboleda, Ryan Arboleda, Sophie Haak, Sam Hjelmeset, Andrew Franck, Bingrui Yang, Jose Bustamante Ortiz, Yuanrong Shen, Joel Walsh

arXiv 2608.10206首次发表:更新:

发表机构

Occidental College(西方学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对儿童语音音素检测难题,提出年龄感知训练方法,构建轻量级模型,开发出可在手机运行的PhonemeTrainer,提升儿童语音ASR与发音辅助应用性能并保障隐私合规。

AI 中文摘要

从儿童语音中检测音素长期以来颇具挑战,原因在于训练数据稀缺以及儿童语音具有独特特征。在一次音素检测竞赛中,我们发现训练一个轻量级模型来同时预测学习者的年龄和音素序列,能让一个拥有9400万参数的模型在目标DrivenData分布上的表现优于拥有3.17亿参数的WavLM Large模型,且与参数规模达其90倍的竞赛集成模型的字符错误率(CER)相差约0.04。这促成了PhonemeTrainer的开发,该应用可在大多数现代手机上运行,最终将为儿童语音提供更优质的自动语音识别(ASR)和发音辅助应用,并带来边缘处理所具备的隐私与合规优势。

英文摘要

Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.

Comments3 pages, 2 figures, 1 table. Demonstration paper presented at the non-archival demonstrations track of the 13th ACM Conference on Learning @ Scale (L@S '26), Seoul, South Korea, June 29-July 3, 2026. Not published in the ACM Digital Library

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑