发表机构
Federal University of Goiás; Federal University of Mato Grosso; São Paulo State University(戈亚斯联邦大学; 马托格罗索联邦大学; 圣保罗州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CVSS-X是一个大规模合成语音到语音翻译语料库,支持英语到28种语言的翻译,总时长超16,000小时,提供标准语音和跨语言克隆两种变体,质量与CVSS相当,促进双向多语种翻译研究。
AI 中文摘要
我们推出了CVSS-X,这是一个大规模合成的语音到语音翻译语料库,通过反转翻译方向扩展了CVSS。CVSS将21种语言翻译成英语,而CVSS-X支持从英语翻译成28种目标语言,涵盖12个语系。该语料库每种语言包含约240,000对平行语音对,总计超过16,000小时,是CVSS的八倍。我们提供两个变体:CVSS-X-C,每种语言有两个标准语音,以及CVSS-X-T,采用跨语言语音克隆,两者均为完全生成。评估显示,其翻译质量与CVSS相当,且在类型多样的语言上表现一致。结合CVSS,这促进了双向和多语种语音到语音翻译的研究。代码可在https URL获取,数据集在CC-BY-NC 4.0许可下于https URL发布。
英文摘要
We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS translates from 21 languages into English, CVSS-X enables translation from English into 28 target languages spanning 12 language families. The corpus comprises approximately 240,000 parallel speech pairs per language, totaling over 16,000 hours, eight times larger than CVSS. We provide two variants: CVSS-X-C with two canonical voices per language, and CVSS-X-T with cross-lingual voice cloning, both fully generated. Evaluation shows comparable translation quality to CVSS with consistent performance across typologically diverse languages. Combined with CVSS, this enables research on bidirectional and multilingual speech-to-speech translation. The code is available at https://github.com/ErmisAI/XVSS-X and the dataset under CC-BY-NC 4.0 license at https://huggingface.co/datasets/lgris/XVSS-X.
CommentsAccepted at the SALMA Workshop (2nd Edition) @ EMNLP 2026 (Non-archival)