发表机构
Aalto University(阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种无需标签或匹配录音的原生参考音素类几何,通过SVD构建坐标系测量L2发音偏差,并与口语熟练度及发音质量显著负相关。
AI 中文摘要
自动口语评估系统能够提供整体熟练度分数,但往往缺乏表征发音质量的可解释度量。我们提出了一种原生参考音素类几何,用于测量第二语言(L2)发音偏差,无需发音标签、朗读提示或来自母语者和L2说话者的相同文本的匹配录音。给定一个母语语音语料库,我们为每个上下文相关的音素类平均帧级自监督表示,并使用奇异值分解(SVD)推导出一个紧凑的原生参考坐标系。对于每个L2话语,我们计算相应的平均值并将其投影到原生参考空间中。然后,我们证明,在Speak and Improve Corpus 2025的Dev子集上,匹配音素类的L2与原生参考坐标之间的距离与整体口语熟练度呈一致的负相关(Spearman's $\ ho\!=\!-0.53$),在English Read by Japanese Students数据集的初学者子集上,与发音质量呈负相关($\ ho\!=\!-0.34$)。这些发现表明,所提出的几何捕获了与熟练度评级相关的声学-语音信息,同时适用于没有匹配原生录音的自发L2语音。
英文摘要
Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's $ρ\!=\!-0.53$) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset ($ρ\!=\!-0.34$). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.
CommentsSubmitted to ICASSP 2027