arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向人类对齐的语音情感相似度评判

Toward Human-Aligned Judgement of Speech Emotion Similarity

Yun-Shao Tsai, Yi-Cheng Lin, Chih-Kai Yang, Ho-Jung Cheng, Tsun-Yi Chang, Sheng-Wei Wu, Yi-Shan Chen, Hsiang-Chun Chang, Liang-Chieh Lee, Hung-yi Lee

arXiv 2609.32504首次发表:更新:

发表机构

National Taiwan University; National Taiwan University of Science and Technology; NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(台湾大学; 台湾科技大学; 台湾大学人工智能卓越研究中心(NTU AI-CoRE))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语音情感相似度评估,提出SES-Bench基准与SES-Judge模型,基于人类比较标注训练,显著优于现有自动度量方法。

AI 中文摘要

在表达性语音生成中评估情感保持涉及衡量生成的语音与参考语音在情感上的接近程度。人类听力测试可评估这种相似性,但其成本促使人们开发与人类判断对齐的自动度量方法。为支持此类度量的开发与评估,我们引入了SES-Bench,一个基于人类对两个候选话语与共享参考话语进行比较而构建的语音情感相似度基准。这些比较记录了听者认为哪个候选话语在情感上更接近参考话语,以及其偏好的强度。利用这些标注,我们训练了SES-Judge来对两个话语之间的情感相似度进行评分。在偏好准确率以及与同时捕捉偏好方向和强度的人类评分的相关性方面,SES-Judge显著优于嵌入余弦相似度和提示式大型音频语言模型。

英文摘要

Evaluating emotion preservation in expressive speech generation involves assessing how closely generated speech matches a reference in emotion. Human listening tests assess this similarity, but their cost motivates automatic measures aligned with human judgments. To support the development and evaluation of such measures, we introduce SES-Bench, a speech emotion similarity benchmark built from human comparisons of two candidate utterances against a shared reference. These comparisons record which candidate listeners find emotionally closer to the reference and the strength of their preference. Using these annotations, we train SES-Judge to score emotion similarity between two utterances. SES-Judge significantly outperforms embedding cosine similarity and prompted large audio-language models in preference accuracy and correlation with human ratings that capture both preference direction and strength.

Comments5 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑