发表机构
Beijing University of Posts and Telecommunications; Beijing DeepLogic Intelligence Technology Co., Ltd.; University of California(北京邮电大学; 北京深逻辑智能科技有限公司; 加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对戏剧配音中文本转语音评估停留在句子级的问题,提出场景级基准SceneTTS-Bench,从音色一致性、情感表现力和节奏连贯性三维度评估,实验表明场景级排名与句子级指标显著不同。
AI 中文摘要
文本转语音系统越来越多地用于戏剧配音,然而评估协议仍停留在句子级别,导致关键的场景级行为未能得到充分衡量。我们提出了SceneTTS-Bench,一个从三个维度评估文本转语音系统的基准:跨角色轮次的音色一致性、高张力话语上的情感表现力,以及分段式长形式合成下的节奏连贯性。该语料库包含真实世界和生成的戏剧脚本,共计160个双语场景和约10,300条话语,其中真实世界脚本作为主要来源(100个场景),生成脚本作为补充来源(60个场景),通过合成数据增强展示了框架的可扩展性。一个与后端无关的规范中间表示确保了跨系统的公平比较。三条自动流水线生成每条话语的诊断结果:用于音色漂移检测的说话人一致性得分、用于识别表演不足的表演不足比率,以及用于量化语速不连续性的语速不连续性比率。在四个文本转语音系统上的实验证实,每个系统都表现出不同的弱点,且场景级排名与句子级指标存在显著差异。基准资源可在以下网址公开获取:此https URL。
英文摘要
Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes and approximately 10,300 utterances, with real-world scripts serving as the primary source (100 scenes) and generated scripts as a supplementary source (60 scenes), demonstrating the framework's extensibility through synthetic data augmentation. A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score for timbre-drift detection, Under-Acting Ratio for under-acting identification, and Rate Discontinuity Ratio for rate-discontinuity quantification. Experiments on four TTS systems confirm that each system exhibits distinct weaknesses and that scene-level rankings diverge substantially from sentence-level metrics. Benchmark resources are publicly available at https://piedpiperg.github.io/scenetts-bench/ .
Comments7 pages, 3 figures; submitted to the ACM Multimedia (ACM MM) 2026 Dataset Track