arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于自监督语音模型的L2发音偏差的母语参考坐标几何方法

A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models

Tina Raissi, Nhan Phan, Mikko Kurimo

arXiv 2609.28060首次发表:更新:

发表机构

Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出母语参考坐标几何方法,利用自监督语音模型将L2发音偏差转化为可解释指标,无需平行录音或发音标签,实验显示与口语熟练度负相关最高达-0.5。

AI 中文摘要

自监督语音模型编码了丰富的语音信息,但如何将这些信息转化为可用于第二语言(L2)自发言语发音评估的可解释指标仍不清楚。我们提出了一种母语参考坐标几何方法,其中来自母语语音的phone类平均值定义了一个低维参考子空间,L2语音通过其与匹配的母语phone类坐标的距离来评估。与先前的基于距离的方法不同,我们的方法不需要具有匹配语言内容的平行录音或专门的发音标签。在不同的自监督编码器和建模选择下,所得的母语参考距离与口语熟练度呈负Spearman相关性,最高可达-0.5,表明熟练度较高的说话者往往更接近母语参考空间。

英文摘要

Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates. Unlike prior distance-based approaches, our method does not require parallel recordings with matched linguistic content or dedicated pronunciation labels. Across different self-supervised encoders and modeling choices, the resulting native-reference distances show negative Spearman correlations up to -0.5 with speaking proficiency, indicating that higher-proficiency speakers tend to lie closer to the native-reference space.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑