arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

YODAS v3:超过100万小时的高带宽、立体声、多语言语音

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe

arXiv 2609.29448首次发表:更新:

发表机构

Carnegie Mellon University; Keio University; Tokyo Metropolitan University; National Institute of Advanced Industrial Science and Technology (AIST)(卡内基梅隆大学; 庆应义塾大学; 东京都立大学; 产业技术综合研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

YODAS v3是一个包含超过110万小时、147种语言、48kHz立体声多声道音频的弱标注语音语料库,通过新技术实现语言平衡收集,并验证了其在语音识别和神经编解码器任务中的有效性。

AI 中文摘要

我们发布了YODAS v3,这是一个弱标注的语音语料库,包含超过110万小时的48kHz多声道音频,涵盖147种语言,以CC BY 3.0许可证发布。YODAS v3不仅是迄今为止最大的开放语音数据集,也是首个真正大规模的高保真立体声语音语料库。我们首先介绍了该语料库的收集方法,其中引入了收集语言平衡语音数据的新技术。通过爬取数据的语言分布证明了我们方法的有效性:YODAS v3中有22种语言的数据超过1万小时,73种语言的数据超过5千小时。随后,我们对数据的构成进行了广泛分析,例如语言分布、音频质量和转录质量。最后,我们训练了基线语音识别和神经编解码器模型,以展示该数据集的有效性。下载请访问此https URL。

英文摘要

We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We first provide the collection methodology for the corpus, where we introduce new techniques for gathering language-balanced speech data. The effectiveness of our approach is shown by the language distribution of the crawled data: 22 languages in YODAS v3 have over 10K hours and 73 languages have over 5K hours of data. We then conduct extensive analyses on the composition of the data, such as the distribution of languages, audio quality, and transcription quality. Finally, we train baseline speech recognition and neural codec models to show the effectiveness of the dataset. Download at https://huggingface.co/datasets/espnet/yodas3.

CommentsInterspeech 2026; 6 Pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑