arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10466eess.AS

音素感知的发音表示用于L2英语母语背景口音识别

Phoneme-Aware Pronunciation Representations for L2-English L1-Background Accent Identification

  • EURECOM(欧洲多媒体通信与网络研究所)

机构由 AI 辅助整理,请以论文原文为准。

Yangyang Qu, Massimiliano Todisco, Nicholas Evans

AI总结:

提出转录辅助模型,用音素级发音单元表示语句,在L2-ARCTIC上实现81.41%准确率,验证音素信息对L2英语口音识别的有效性。

AI中文摘要:

我们研究L2英语中说话者不相交的口音识别,其目标是从英语发音预测说话者的第一语言(L1)背景。大多数现有系统使用单一的语句级表示对口音进行分类,但这种全局表示可能掩盖依赖于特定英语音素的发音线索。我们提出了一种转录辅助模型,在口音识别过程中显式地利用音素信息。我们不是仅将语句表示为全局语音嵌入,而是将其表示为一系列发音单元,每个单元将来自口语片段的声学证据与该片段对齐的英语音素相结合。冻结的语音编码器提供声学特征,而转录仅用于获取音素级强制对齐。没有词级或句子级文本表示传递给口音分类器。在L2-ARCTIC上的四折说话者不相交协议下,我们的模型达到了81.41%的准确率和81.21%的宏F1分数,这是所评估系统中平均性能最高的。诊断性消融实验支持了音素对齐标记构建的重要性,而基于Whisper的消融实验显示了音素信息带来的额外收益。

英文摘要:

We study speaker-disjoint accent identification for L2 English, where the goal is to predict a speaker's first-language (L1) background from English pronunciation. Most existing systems classify accents using a single utterance-level representation, but such global representations can obscure pronunciation cues that depend on specific English phonemes. We propose a transcript-assisted model that makes phoneme information explicit during accent identification. Instead of representing an utterance only as a global speech embedding, we represent it as a sequence of pronunciation units, each combining acoustic evidence from a spoken segment with the aligned English phoneme for that segment. A frozen speech encoder provides the acoustic features, while the transcript is used only to obtain phoneme-level forced alignments. No word-level or sentence-level text representation is passed to the accent classifier. Under a four-fold speaker-disjoint protocol on L2-ARCTIC, our model achieves 81.41% accuracy and 81.21% macro-F1, the highest mean performance among the evaluated systems. Diagnostic ablations support the importance of phoneme-aligned token construction, while a Whisper-based ablation shows an additional gain from phoneme information.

↑