arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SceneTTS-Bench:戏剧配音中的场景级文本转语音基准

SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing

Yizhong Geng, Yanliang Li, Jinghan Yang, Tianhan Jiang, Yingming Gao, Ya Li

arXiv 2609.26255首次发表:更新:

发表机构

Beijing University of Posts and Telecommunications; Beijing DeepLogic Intelligence Technology Co., Ltd.; University of California(北京邮电大学; 北京深逻辑智能科技有限公司; 加利福尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对戏剧配音中文本转语音评估停留在句子级的问题,提出场景级基准SceneTTS-Bench,从音色一致性、情感表现力和节奏连贯性三维度评估,实验表明场景级排名与句子级指标显著不同。

AI 中文摘要

文本转语音系统越来越多地用于戏剧配音,然而评估协议仍停留在句子级别,导致关键的场景级行为未能得到充分衡量。我们提出了SceneTTS-Bench,一个从三个维度评估文本转语音系统的基准:跨角色轮次的音色一致性、高张力话语上的情感表现力,以及分段式长形式合成下的节奏连贯性。该语料库包含真实世界和生成的戏剧脚本,共计160个双语场景和约10,300条话语,其中真实世界脚本作为主要来源(100个场景),生成脚本作为补充来源(60个场景),通过合成数据增强展示了框架的可扩展性。一个与后端无关的规范中间表示确保了跨系统的公平比较。三条自动流水线生成每条话语的诊断结果:用于音色漂移检测的说话人一致性得分、用于识别表演不足的表演不足比率,以及用于量化语速不连续性的语速不连续性比率。在四个文本转语音系统上的实验证实,每个系统都表现出不同的弱点,且场景级排名与句子级指标存在显著差异。基准资源可在以下网址公开获取:此https URL。

英文摘要

Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes and approximately 10,300 utterances, with real-world scripts serving as the primary source (100 scenes) and generated scripts as a supplementary source (60 scenes), demonstrating the framework's extensibility through synthetic data augmentation. A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score for timbre-drift detection, Under-Acting Ratio for under-acting identification, and Rate Discontinuity Ratio for rate-discontinuity quantification. Experiments on four TTS systems confirm that each system exhibits distinct weaknesses and that scene-level rankings diverge substantially from sentence-level metrics. Benchmark resources are publicly available at https://piedpiperg.github.io/scenetts-bench/ .

Comments7 pages, 3 figures; submitted to the ACM Multimedia (ACM MM) 2026 Dataset Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑