自动语音识别中的中文幽默同音字识别与歧义消解
Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition
AI总结:
该研究针对第二语言中文发音学习的MDD需求,提出基于音系特征的Wav2Vec2-CTC框架,区分音段与声调错误,可降低FAR与DER,提供更详细的诊断反馈。
AI中文摘要:
自动发音错误检测与诊断(MDD)在第二语言中文发音学习中至关重要。尽管基于端到端(E2E)的MDD方法已大幅提升音素级检测准确率,但诊断反馈仍有限,因为未明确区分音段与声调错误。本文提出一种基于音系特征的MDD框架,在统一的Wav2Vec2-CTC架构中同时建模音段与声调属性。实验结果显示,与仅采用音素的基线系统相比,该方法将误接受率(FAR)降低10.1%,诊断错误率(DER)降低23.6%。通过将音素分解为低级音系组件,该方法能为第二语言学习者提供更详细、可解释的诊断反馈。
英文摘要:
Mandarin homophones remain a key challenge to improving automatic speech recognition (ASR) accuracy due to the amount of potential homophones. Mandarin speakers use this feature casually to convey emotions such as humour. Recent homophone-aware ASR studies have improved recognition accuracy, but intentional homophone twists in speech remain underexplored. In this paper, we identify patterns of homophone-based rhetorical wordplay in Mandarin, referred to as HumourPhone, and propose an ASR Adapter for homophone and HumourPhone recovery. Experimental results show that the proposed approach improves the recall of recognising HumourPhone by over 5\% and achieves a 4.35\% drop in target-span character error rate for homophone correction compared to baseline. These results highlight the need for homophone-aware modelling of lexical ambiguity and rhetorical wordplay in Mandarin speech recognition.