发表机构
Northwestern Polytechnical University; Nanjing University; University of New South Wales; WeNet Open Source Community; Moonstep AI; Nexdata; Xiamen University; State Key Laboratory of Novel Software Technology, Nanjing University(西北工业大学; 南京大学; 新南威尔士大学; WeNet开源社区; Moonstep AI; Nexdata; 厦门大学; 南京大学软件新技术国家重点实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对闽南语语音资源稀缺问题,构建约1万小时带闽南语与普通话双转录的WenetSpeech-Min语料库,并建立ASR与TTS基准,训练模型在多数指标上优于开源系统且接近商业水平。
AI 中文摘要
方言语音技术的进展受到大规模真实世界语料库稀缺的阻碍。对于闽南语语音,现有资源仍然有限,且很少有资源提供大规模配对的闽南语和普通话转录。为解决这些空白,我们推出了WenetSpeech-Min,一个开源语料库,包含从多样在线媒体收集的约10,000小时闽南语语音,并为每个话语提供配对的闽南语和普通话转录。我们进一步建立了一个自动语音识别(ASR)基准,涵盖闽南语和普通话转录,以及一个使用闽南语转录的文本到语音合成(TTS)基准,并为这两项任务提供了人工验证的评估集。为评估语料库的有效性,我们在WenetSpeech-Min上训练ASR和TTS模型,并将其与提议基准上的代表性系统进行比较。所得模型在大多数指标上优于评估的开源模型,并达到与商业系统竞争的性能。我们将发布语料库、基准和模型,以促进闽南语语音技术的可复现研究。
英文摘要
Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.
Comments5 pages, 2 figures