VietPrism:包含多样方言与语码转换的大规模越南语语音与深度伪造语料库
VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
浏览论文内容
中文总结 AI 辅助
VietPrism是首个大规模越南语多领域语料库,整合了转录、说话人身份、五种方言及自然语码转换,并包含受控生成的伪造语音,用于评估多语言深度伪造检测器的脆弱性。
中文摘要 AI 辅助
越南语语音研究受到资源限制,这些资源将自动语音识别与说话人、方言、语码转换和深度伪造分析分离开来。我们引入了VietPrism,一个开放的多领域语料库,将这些维度大规模地结合在一起:包含来自8,388个真实世界视频的1,262位经过验证的说话人的993.4小时和403,941条真实语音话语。据我们所知,这是第一个大规模越南语语料库,同时提供转录文本、一致的说话人身份、五种方言组以及自然发生的越南语-英语语码转换,按时长计算,语码转换几乎占语料库的一半。我们进一步使用四种开源和商业合成系统创建了超过3.1K小时的伪造语音。每个伪造样本都以经过验证的说话人参考为条件,并与转录和说话人匹配的真实语音话语配对,从而实现独特的受控评估,减少词汇和身份混淆。对五个预训练多语言检测器的零样本评估揭示了惊人的脆弱性:等错误率(EER)在检测器-生成器配对之间差异很大,而最近的多语言检测器DFA-1B在说话人相似度增加时,其EER从16.3%恶化到33.6%。按方言分层的结果进一步暴露了模型相关的差异。通过将自然语言多样性与受控伪造生成相结合,VietPrism为越南语语音建模和可信音频深度伪造检测提供了一个具有挑战性的基础。
英文摘要
Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.
发表机构
- Independent Researcher(独立研究者)
- Indiana University, Bloomington, USA(印第安纳大学布鲁明顿分校)
机构由 AI 辅助整理,请以论文原文为准。