发表机构
NC AI Co., Ltd, Republic of Korea; Sogang University, Republic of Korea(韩国NC AI公司; 韩国成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对设计发声缺乏公开资源及研究不足的问题,构建设计发声数据集,通过整合多样原始发声源并处理生成变体,提供标准化测试集,报告基线基准结果,以支持语音转换相关评估与研究。
AI 中文摘要
基于人工智能的语音转换进展推动了包括电影、有声读物和游戏在内的广泛媒体应用。然而,大多数研究和公共基准仍聚焦于自然人类语音,设计发声(如怪物咆哮和机器人声音)未得到充分探索,部分原因是缺乏公开可用资源。为填补这一空白,我们引入设计发声数据集,通过策划包括语音和动物发声在内的各种原始发声源,并应用专业语音效果处理来生成相应的效果修改变体。我们还提供了一个标准化测试集,在源音色组和预设风格上有明确的可见/不可见分割,以评估受控条件下的泛化能力。最后,我们报告基线基准结果以支持可重复评估和未来研究。数据集和演示样本可通过此https网址获取。
英文摘要
Advances in AI-based voice conversion have enabled a wide range of media applications, including films, audiobooks, and games. However, most research and public benchmarks still focus on natural human speech, leaving designed vocalizations, such as monster growls and robotic voices, underexplored, partly due to the lack of publicly available resources. To address this gap, we introduce the Designed Vocalizations Dataset, constructed by curating diverse raw vocal sources, including speech and animal vocalizations, and applying professional vocal effects processing to produce corresponding effect modified variants. We further provide a standardized test set with explicit seen/unseen splits over source timbre groups and preset styles to assess generalization under controlled conditions. Finally, we report baseline benchmark results to support reproducible evaluation and future research. The dataset and demo samples are available at https://ncai-official.github.io/speech/publications/designed-vocalizations-dataset/.
CommentsAccepted at InterSpeech 2026