arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WenetSpeech-Min:用于方言语音处理的大规模闽南语语音语料库及双转录

WenetSpeech-Min: A Large-Scale Minnan Speech Corpus with Dual Transcriptions for Dialectal Speech Processing

Haoyu Zhang, Chunjiang He, Hongtao Li, Zeyu Zhu, Qituan Shangguan, Chengyou Wang, Jingbin Hu, Ziyu Zhang, Bingshen Mu, Yanbo Wang, Shuai Wang, Jinhui Ye, Chengdong Liang, Binbin Zhang, Pengcheng Zhu, Chuang Ding, Qianze Feng, Qingyang Hong, Liumeng Xue, Lei Xie

arXiv 2609.36834首次发表:更新:

发表机构

Northwestern Polytechnical University; Nanjing University; University of New South Wales; WeNet Open Source Community; Moonstep AI; Nexdata; Xiamen University; State Key Laboratory of Novel Software Technology, Nanjing University(西北工业大学; 南京大学; 新南威尔士大学; WeNet开源社区; Moonstep AI; Nexdata; 厦门大学; 南京大学软件新技术国家重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对闽南语语音资源稀缺问题,构建约1万小时带闽南语与普通话双转录的WenetSpeech-Min语料库,并建立ASR与TTS基准,训练模型在多数指标上优于开源系统且接近商业水平。

AI 中文摘要

方言语音技术的进展受到大规模真实世界语料库稀缺的阻碍。对于闽南语语音,现有资源仍然有限,且很少有资源提供大规模配对的闽南语和普通话转录。为解决这些空白,我们推出了WenetSpeech-Min,一个开源语料库,包含从多样在线媒体收集的约10,000小时闽南语语音,并为每个话语提供配对的闽南语和普通话转录。我们进一步建立了一个自动语音识别(ASR)基准,涵盖闽南语和普通话转录,以及一个使用闽南语转录的文本到语音合成(TTS)基准,并为这两项任务提供了人工验证的评估集。为评估语料库的有效性,我们在WenetSpeech-Min上训练ASR和TTS模型,并将其与提议基准上的代表性系统进行比较。所得模型在大多数指标上优于评估的开源模型,并达到与商业系统竞争的性能。我们将发布语料库、基准和模型,以促进闽南语语音技术的可复现研究。

英文摘要

Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.

Comments5 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑