用于第二语言语音、节奏和语调评分的自监督语音比较
Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
浏览论文内容
中文总结 AI 辅助
研究探索基于自监督WavLM表示的DTW能否为二语语音评估提供无文本框架,结果显示该方法在语音评分上超人类一致性,在节奏评估上接近人类水平,在语调评估上部分任务表现一般,表明自监督表示是多方面发音评估的有前景基础。
中文摘要 AI 辅助
传统上,第二语言语音评估侧重于语音评估,对节奏和语调等超音段特征的评分探索不足。此外,评估方法通常需要用带标签的第二语言语音数据进行训练,难以应用于低资源环境。我们研究基于自监督WavLM表示的动态时间规整(DTW)是否能为评估英日第二语言语音的语音准确性、节奏和语调提供无文本框架。结果表明,将学习者语音与母语模板进行比较的基于基本DTW的方法在整体和句子级语音评分上超过了人类一致性。对于节奏,我们引入了测量DTW对齐路径中扭曲程度的方法,最佳方法接近人类水平。对于语调,我们将韵律残差上的DTW距离与音高和强度特征相结合,但在某些任务上性能仍较为一般。我们的结果表明自监督表示是多方面发音评估的有前景的无文本基础。
英文摘要
L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. Moreover, assessment methods often require training with labeled L2 speech data, making them difficult to apply in low-resource settings. We investigate whether DTW over self-supervised WavLM representations can provide a text-free framework for assessing phonetic accuracy, rhythm, and intonation in English and Japanese L2 speech. Results show that a basic DTW-based approach that compares learner speech to native templates exceeds human agreement on holistic and sentence-level phonetic scoring. For rhythm, we introduce methods that measure the degree of warping in the DTW alignment path; our best method approaches human-level performance. For intonation, we combine DTW distance over prosodic residuals with pitch and intensity features, but performance remains more modest on some tasks. Our results point to self-supervised representations as a promising, text-free basis for multi-aspect pronunciation assessment.