arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13792cs.SDeess.AS

VoiceMOS 挑战赛2026:评估语音增强、情感文本转语音和带口音文本转语音系统

The VoiceMOS Challenge 2026: Evaluating Speech Enhancement, Emotional TTS and Accented TTS Systems

Wen-Chin Huang, Wei Wang, Marvin Sach, Xiaoxue Gao, Nicholas Sanders, Erica Cooper, Toda Tomoki

首次发表
浏览论文内容

中文总结 AI 辅助

VoiceMOS挑战赛2026评估语音增强、情感TTS和带口音TTS系统,设三个赛道,吸引18支队伍,多数超越基线,并总结结果与未来方向。

中文摘要 AI 辅助

我们展示了VoiceMOS挑战赛2026的结果,这是关于主观语音评估自动预测的第五版科学挑战赛。在2025年将范围扩展到音乐和通用音频后,我们重新将评估目标聚焦于语音,并组织了三个赛道:预测增强语音的绝对和比较类别评级,预测情感文本转语音系统的自然度和情感相关任务,以及预测基于编解码器的语音合成系统的说话人和口音相似度。该挑战赛吸引了全球共18支队伍,大多数队伍成功超越了提供的基线。我们总结了挑战赛结果、具有代表性的顶尖系统、参与者反馈以及未来版本的改进方向。

英文摘要

We present the results of the VoiceMOS Challenge 2026, the fifth edition of a scientific challenge on automatic prediction of subjective speech assessments. After expanding the scope to music and general audio in 2025, we refocused the evaluation target on speech and organized three tracks: prediction of absolute and comparative category ratings for enhanced speech, prediction of naturalness and emotion-related tasks for emotional text-to-speech systems, and prediction of speaker and accent similarity for codec-based speech synthesis systems. The challenge attracted a total of 18 teams worldwide, with most teams successfully surpassing the provided baselines. We summarize the challenge results, representative top-performing systems, participant feedback, and directions for future editions.

发表机构

  • Nagoya University(名古屋大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • Technische Universität Braunschweig(布伦瑞克工业大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • University of Edinburgh(爱丁堡大学)
  • National Institute of Information and Communications Technology(日本国立信息通信技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑