PersianVox:一种面向野外数据的语音数据集生成的韵律感知方法
PersianVox: A Prosody-Aware Approach for Speech Dataset Generation from In-the-Wild Data
浏览论文内容
中文总结 AI 辅助
PersianVox提出一种韵律感知的全自动流水线,利用声学话轮检测和双模型一致性机制,从野外数据生成2,400小时波斯语多说话人语音数据集,并首次提供波斯语语音质量评估基准。
中文摘要 AI 辅助
目前,零样本文本到语音合成的发展在低资源语言中受到大规模、高保真语音数据集稀缺的阻碍。传统的基于对齐的方法需要罕见的逐字转录文本,而标准的野外数据流水线通常依赖单模型自动语音识别和基于静音的切分,导致转录错误和韵律截断。为了解决波斯语中的这些挑战,本文介绍了PersianVox,一个全自动流水线,旨在从未标记的网络数据中生成高质量语音语料库。我们的方法整合了一种新颖的韵律感知切分策略,该策略利用声学话轮检测来保持语言完整性,并优化话语时长以适应长上下文建模。此外,我们采用双模型一致性机制,利用两种不同的模型架构来过滤不可靠的转录,而无需真实标签。该流水线生成了一个2,400小时的多说话人数据集,这是迄今为止波斯语可用的最大开源语音资源。此外,我们提供了首个波斯语语音质量评估方法的比较基准,并发布了一个人工标注的子集以促进未来研究。
英文摘要
Advancement of zero-shot text-to-speech synthesis is currently hindered for low-resource languages by the scarcity of large-scale, high-fidelity speech datasets. Traditional alignment-based methods require rare verbatim transcripts, while standard in-the-wild pipelines often rely on single-model automatic speech recognition and silence-based segmentation, leading to transcription errors and truncated prosody. To address these challenges for the Persian language, this paper introduces PersianVox, a fully automated pipeline designed to generate high-quality speech corpora from unlabeled web data. Our approach integrates a novel prosody-aware segmentation strategy that utilizes acoustic turn-detection to preserve linguistic completeness and optimize utterance duration for long-context modeling. Furthermore, we employ a dual-model agreement mechanism, leveraging two distinct model architectures to filter unreliable transcriptions without ground truth. This pipeline yields a 2,400-hour multi-speaker dataset, the largest open-source speech resource available for Persian to date. Additionally, we provide the first comparative benchmark of speech quality assessment methods for Persian, releasing a human-annotated subset to facilitate future research.
发表机构
- Electronics Research Institute, Sharif University of Technology(电子研究所,谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。