通过合成视觉数据增强改进音视频语音识别
Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation
浏览论文内容
中文总结 AI 辅助
本研究利用音频驱动的说话人头像生成合成视觉数据,作为增强和独立训练资源,在西班牙语和加泰罗尼亚语上实现最高16.2%的相对词错误率降低,证明其可扩展解决AVSR数据稀缺问题。
中文摘要 AI 辅助
音视频语音识别(AVSR)是一种多模态语音识别方法,它结合了来自唇部运动的视觉信息以提升模型性能。尽管具有这些优势,其发展仍受到标记音视频(AV)数据集有限性的制约。本研究探索使用合成视觉数据作为解决方案,利用音频驱动的说话人头像流程从现有音频数据生成唇部同步的视觉内容。我们评估了合成视觉数据作为增强策略和独立训练资源的有效性,并将其应用于西班牙语和加泰罗尼亚语。我们的结果表明,用合成样本增强真实音视频数据可带来相对词错误率(WER)最高达16.2%的降低,展示了该方法的潜力。此外,我们证明合成数据单独即可作为缺乏音视频数据集的语言中AVSR训练的基线。这些发现提供了证据,表明合成视觉数据可作为AVSR数据稀缺性的可扩展解决方案,实现更广泛的语言覆盖。
英文摘要
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. This work explores the use of synthetic visual data as a solution, using an audio-driven talking-head pipeline to generate lip-synchronized visual content from existing audio data. We evaluate the effectiveness of synthetic visual data both as an augmentation strategy and as a standalone training resource, applying our approach to Spanish and Catalan. Our results show that augmenting real AV data with synthetic samples yields relative Word Error Rate (WER) reductions of up to 16.2%, demonstrating the potential of this approach. Moreover, we demonstrate that synthetic data alone can serve as a baseline for AVSR training in languages lacking AV datasets. These findings provide evidence that synthetic visual data can serve as a scalable solution to AVSR data scarcity, enabling broader language coverage.
发表机构
- Universitat Politècnica de Catalunya (UPC)(加泰罗尼亚理工大学)
- Barcelona Supercomputing Center (BSC)(巴塞罗那超级计算中心)
机构由 AI 辅助整理,请以论文原文为准。