arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09930cs.SDcs.AIcs.CL

超越自然性:基于语言维度探究自动文本转语音评估器

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建了首个TTS维度级元评估基准,测试发现MOS预测器聚焦声学信号质量,Audio-LLM评估器检测能力具选择性且依赖提示,两类方法均难捕捉语言结构化语音错误,相关资源已公开。

中文摘要 AI 辅助

自动文本转语音(TTS)评估方法(包括平均意见得分(MOS)预测器和音频大语言模型(Audio-LLM)评估器)应反映人类感知,但目前尚不清楚它们在多大程度上捕捉到听者实际感知的不同语音方面。我们将“自然性”解构为一个基于语言的标注架构,涵盖10个不同的感知维度,并以此构建首个TTS的维度级元评估基准,该基准包含由受过训练的语言学家标注员标注的860个话语。对4个MOS预测器和4个Audio-LLM评估器的基准测试结果显示,MOS预测器聚焦于声学信号质量,而Audio-LLM评估器表现出选择性的、依赖提示的检测能力,无法在所有维度上泛化。两类方法均无法可靠捕捉广泛的语言结构化语音错误。我们的数据集、标注架构和评估代码已公开发布,以支持更具针对性和可解释性的TTS评估。

英文摘要

Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.

发表机构

  • ServiceNow(ServiceNow公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑