arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33060cs.LG

深度学习技术在意大利儿童语音音素识别中的应用

Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech

  • University of Turin(都灵大学)
  • Fondazione Paideia Ente Filantropico(派迪亚慈善基金会)

机构由 AI 辅助整理,请以论文原文为准。

Nicola Barbaro, Cristina Gena, Francesco Petriglia, Andrea Meirone, Alessandro Mazzei, Arianna Viotti

AI总结:

本研究提出基于Conformer的Broca系统,利用成人语音预训练和少量儿童语音微调,在意大利儿童语音音素识别中达到13.36%的错误率,证明低资源下实现高效语音到IPA转录的可行性。

AI中文摘要:

言语治疗师在诊断言语障碍时,常因缺乏将语音转录为国际音标(IPA)的高效工具而面临困难。本研究通过Broca系统应对这一挑战,该系统是一个基于Conformer的深度学习系统,在8天的成人语音上进行了预训练,并在一个165分钟的意大利儿童语音数据集上进行了微调,该数据集通过一系列针对3.5至6.5岁儿童的标准化诊断测试收集。Broca经过优化以处理儿童语音中的语音变异性,包括语调、口音和言语错误,并在意大利语音上实现了最先进的加权音素错误率13.36%。值得注意的是,这一性能是在使用不到三小时的儿童特定数据的情况下获得的,凸显了该模型在低资源临床环境中的效率和鲁棒性。这项工作表明,利用最少的数据即可实现准确的、与词汇无关的语音到IPA转录,为开发更易获取、数据高效的言语评估与诊断支持工具铺平了道路。

英文摘要:

Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-minute dataset of Italian child speech collected through a range of standardized diagnostic tests for children aged 3.5-6.5. Broca was optimized to handle phonetic variability in children's speech, including tone, accent, and speech errors, and achieved a state-of-the-art weighted Phoneme Error Rate of 13.36% on Italian speech. Remarkably, this performance was obtained using less than three hours of child-specific data, underscoring the model's efficiency and robustness in low-resource clinical settings. This work demonstrates that accurate, vocabulary-independent speech-to-IPA transcription can be achieved with minimal data, paving the way for more accessible, data-efficient tools to support speech assessment and diagnosis.

↑