arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14846cs.SDcs.AI

RW-Voice-EQ基准:评估语音人工智能系统的真实世界基准

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

  • Hume AI Research(休谟人工智能研究)

机构由 AI 辅助整理,请以论文原文为准。

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghba… 展开作者

David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis

AI总结:

研究提出真实世界语音EQ基准,用于跨TTS、STS、SU和ASR评估语音人工智能,指出当前基准多评估孤立能力,新基准能考量声学等信息,评估表明性能依赖维度,语音人工智能应综合多方面能力评估而非单一总分。

AI中文摘要:

当前语音人工智能基准通常评估孤立能力,如语音清晰度、单词错误率或基于文本的对话质量,很少测试系统是否利用区分口语与其文本表示的声学信息。为此,我们引入了真实世界语音EQ基准,这是一个用于跨文本到语音(TTS)、语音到语音(STS)、语音理解(SU)和自动语音识别(ASR)评估语音人工智能的多维基准。我们的评估表明性能高度依赖维度。对于TTS,自然度、表现力、身份稳定性和可靠性在很大程度上是独立的评估维度。对于STS,访问音频并不保证使用声音情感,一些智能体在很大程度上仍由转录驱动。对于SU,模型在跨副语言任务中的表现参差不齐。对于ASR,现实世界中的口音、情感、噪音和对话条件暴露了既定纯净语音基准未捕捉到的失败。这些结果共同表明,语音人工智能应作为声学、表现力、交互性和鲁棒性能力的概况进行评估,而不是通过单一总分。

英文摘要:

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

补充信息

↑