对多样化语音的人类与自动语音识别进行基准测试:初步结果
Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
浏览论文内容
中文总结 AI 辅助
研究对多样化语音(荷兰儿童、老年人语音及弗拉芒语)中人类与自动语音识别性能进行初步比较,发现谷歌电话识别系统表现优,ASR系统与听众表现相似,特定情况超越听众,还发现相关性能差异,未来研究应增强ASR系统对声学变异性的鲁棒性。
中文摘要 AI 辅助
人类常被视为最佳听众及自动语音识别(ASR)系统的性能上限。本文对最先进的ASR系统和荷兰本土听众在识别“多样化”语音(特别是荷兰儿童和老年人语音以及弗拉芒语)方面的性能进行了初步比较。谷歌电话识别系统表现优于其他ASR系统。重要的是,ASR系统与听众表现相似,特定情况下甚至超越听众。发现听众与ASR系统在说话者年龄、地区口音和话语长度方面存在细微性能差异。未来研究应聚焦使ASR系统对与衰老和地区口音相关的声学变异性更具鲁棒性。对测试刺激和完整Jasmin - CGN测试集上的ASR识别性能比较显示了特定测试集对人类和ASR性能基准测试结论的影响。
英文摘要
Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of "diverse" speech, specifically Dutch child and older adults' speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.