基于课程学习的噪声自适应方法用于视觉语音识别中的音素到文本重建
Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition
- Kyushu Institute of Technology(九州工业大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视觉语音识别中音素到文本重建模型训练与推理不匹配的问题,提出渐进式错误课程训练(PECT),通过逐步引入真实音素错误提升鲁棒性,在LRS2和LRS3上显著降低词错误率。
AI中文摘要:
以音素为中心的视觉语音识别从中间音素预测中重建句子,使得整体识别性能高度依赖于音素到文本重建模型的鲁棒性。现有的重建方法通常是在干净的音素序列或合成损坏的输入上训练的,导致训练条件与推理过程中遇到的实际音素预测错误之间存在不匹配。为了解决这一局限性,本文提出了渐进式错误课程训练(PECT)。这一课程学习框架使用合成音素扰动、多域伪标签以及由视觉语音识别器生成的目标域伪标签,逐步适应基于“不落下任何语言”(NLLB)的音素到文本重建模型。通过逐渐让重建模型接触越来越真实的音素预测错误,所提出的框架在保持句子重建准确性的同时提高了鲁棒性。在LRS2和LRS3基准上的实验表明,PECT在多种基于音素的视觉语音识别前端上持续改善重建性能,这些前端包括视觉自动语音识别(V-ASR)、点视觉自动语音识别(PV-ASR)和头部姿态感知视觉语音识别(HP-VSR)变体。特别是,PECT将HP-VSR-FiLMFuse(L4)在LRS2上的词错误率(WER)从23.3%降低到22.2%,并将HP-VSR-ResFiLM在LRS3上的WER从30.3%降低到29.7%。全面的消融研究和定性分析进一步证明了将重建模型逐步适应真实音素预测错误的有效性。这些结果表明,PECT为以音素为中心的视觉语音识别中的音素到文本重建提供了一种有效且可推广的课程学习策略。
英文摘要:
Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.