arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dialogs:用于对话助手的具有工作室质量的富有表现力的俄语对话语料库

Dialogs: a studio-quality expressive conversational Russian speech corpus for dialog assistants

Ilya Shigabeev, Ilya Latyshev

arXiv 2607.14310首次发表:更新:

AI 中文总结

介绍用于对话助手的俄语对话语料库Dialogs,含20.6小时专业录制对话及风格情感标签。经人群MOS测试验证质量,虽每个说话者数据有限,但支持训练富有表现力、类似对话的TTS,还以VITS2模型作概念验证。

AI 中文摘要

我们介绍了Dialogs,一个用于对话助手的具有工作室质量的俄语对话语音语料库。该数据集包含在专业工作室录制的20.6小时面对面表演对话(44.1kHz立体声),分为3个说话者的11796个话语。与朗读语音资源不同,Dialogs捕捉了轮流节奏和富有表现力的韵律,并提供了涵盖12类别的每个话语风格/情感标签。我们通过人群MOS测试验证了语料库质量,显示出与强大的俄语工作室基线相当的音频质量和清晰度,同时在表现力和对话自然度方面获得更高评分。最后,我们训练了一个VITS2模型作为概念验证,证明Dialogs支持训练富有表现力、类似对话的TTS,尽管每个说话者的数据有限。

英文摘要

We introduce Dialogs, a studio-quality Russian conversational speech corpus for dialog assistants. The dataset contains 20.6 hours of face-to-face acted dialogs recorded in a professional studio (44.1 kHz stereo) and segmented into 11,796 utterances across 3 speakers. Unlike read-speech resources, Dialogs captures turn-taking rhythm and expressive prosody, and provides per-utterance style/emotion labels spanning 12 categories. We validate corpus quality with crowd MOS tests, showing comparable audio quality and intelligibility to strong Russian studio baselines while achieving higher ratings for expressiveness and conversational naturalness. Finally, we train a VITS2 model as a proof of concept, demonstrating that Dialogs supports training expressive, dialog-like TTS despite limited per-speaker data.

Comments4 pages, 1 figure, 5 tables. Interspeech 2026

DOI:10.57967/hf/9194

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑