发表机构
King Hussein School of Computing Sciences(侯赛因国王计算科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一个端到端基于Transformer的约旦方言语音转文本自监督学习框架,利用噪声学生训练和自训练,在词错误率上比微调Wav2Vec降低5%,并提供数据集以支持低资源语言应用。
AI 中文摘要
语音到文本引擎在当今的各种应用中极为需要,是人机交互中的关键使能技术。然而,一些语言,尤其是阿拉伯方言或任何低资源语言,缺乏带标签的语音数据。自监督训练过程和使用噪声训练的自训练已被证明是新兴的可行解决方案之一。本文提出了一种基于Transformer的端到端模型及框架,适用于低资源语言。此外,该框架整合了定制的音频到文本处理算法,以实现高效的约旦阿拉伯方言语音到文本系统。所提出的框架能够从多个来源摄取数据,通过加速手动标注过程,使得从外部来源获取真实标签成为可能。该框架允许使用噪声学生训练和自监督学习进行训练,以在预训练和后训练阶段利用未标注数据,并整合多种类型的数据增强。所提出的自训练方法在词错误率降低方面比微调的Wav2Vec模型高出5%。这项工作的成果为研究社区提供了一个约旦口语数据集以及一种处理低资源语言的端到端方法。这是通过利用预训练、后训练的力量,并以最少的人工干预注入带噪声的标注数据和增强数据来实现的。它使得在阿拉伯语语音到文本领域开发新应用成为可能,如问答系统和智能控制系统,并将为智能机器人增添类似人类的感知和听觉传感器。
英文摘要
Speech-to-text engines are extremely needed nowadays for different applications, representing an essential enabler in human-robot interaction. Still, some languages suffer from the lack of labeled speech data, especially in the Arabic dialects or any low-resource languages. The need for a self-supervised training process and self-training using noisy training is proven to be one of the up-and-coming feasible solutions. This article proposes an end-to-end, transformers-based model with a framework for low-resource languages. In addition, the framework incorporates customized audio-to-text processing algorithms to achieve a highly efficient Jordanian Arabic dialect speech-to-text system. The proposed framework enables ingesting data from many sources, making the ground truth from external sources possible by speeding up the manual annotation process. The framework allows the training process using noisy student training and self-supervised learning to utilize the unlabeled data in both pre- and post-training stages and incorporate multiple types of data augmentation. The proposed self-training approach outperforms the fine-tuned Wav2Vec model by 5% in terms of word error rate reduction. The outcome of this work provides the research community with a Jordanian-spoken data set along with an end-to-end approach to deal with low-resource languages. This is done by utilizing the power of the pretraining, post-training, and injecting noisy labeled and augmented data with minimal human intervention. It enables the development of new applications in the field of Arabic language speech-to-text area like the question-answering systems and intelligent control systems, and it will add human-like perception and hearing sensors to intelligent robots.
Journal refFront. Robot. AI 9:1090012 (2022)
DOI:10.3389/frobt.2022.1090012