arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越一刀切:通过个体情感感知空间的交互式优化实现个性化与文化自适应的情感文本转语音(TTS)

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

Wangzixi Zhou, Bagus Tris Atmaja, Sakriani Sakti

arXiv 2608.00998首次发表:更新:

AI 中文总结

该研究针对现有情感TTS一刀切的缺陷,提出采用交互式遗传算法优化个体唤醒-效价感知空间的个性化文化自适应情感TTS框架,经多文化参与者评估验证其有效性。

AI 中文摘要

对话式AI的兴起提升了对情感文本转语音(TTS)的关注度,多数系统依赖离散情感标签,无法捕捉人类情感的细微差别,近期模型采用罗素的唤醒-效价(A-V)模型等维度表征,提供更精细的控制,但个体与文化间的情感感知存在差异,可能导致建模情感与感知情感不匹配。我们提出一种个性化与文化自适应的情感TTS框架,采用交互式遗传算法对个体A-V感知空间进行交互式优化,通过调整每个听者的情感表征,系统生成的语音比使用平均A-V值的模型具有更符合感知的情感表达,对日本、中国、印尼参与者的评估凸显了个性化与文化自适应对超越一刀切情感TTS的重要性。

英文摘要

The rise of conversational AI has increased interest in emotional Text-to-Speech (TTS). Most systems rely on discrete emotion labels, which fail to capture the nuanced nature of human affect. Recent models employ dimensional representations such as Russell's arousal-valence (A-V) model, offering finer control. However, emotional perception varies across individuals and cultures, which may cause mismatches between modeled and perceived emotions. We propose a personalized and culturally adaptive emotional TTS framework that performs interactive optimization of individualized A-V perception spaces using an Interactive Genetic Algorithm. By adapting emotion representations to each listener, the system produces speech with more perceptually aligned emotional expression than models using averaged A-V values. Evaluations with Japanese, Chinese, and Indonesian participants highlight the importance of personalization and cultural adaptation for moving beyond one-size-fits-all emotional TTS.

CommentsAccepted for publication at INTERSPEECH 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑