AI 中文总结
该研究针对现有情感TTS一刀切的缺陷,提出采用交互式遗传算法优化个体唤醒-效价感知空间的个性化文化自适应情感TTS框架,经多文化参与者评估验证其有效性。
AI 中文摘要
对话式AI的兴起提升了对情感文本转语音(TTS)的关注度,多数系统依赖离散情感标签,无法捕捉人类情感的细微差别,近期模型采用罗素的唤醒-效价(A-V)模型等维度表征,提供更精细的控制,但个体与文化间的情感感知存在差异,可能导致建模情感与感知情感不匹配。我们提出一种个性化与文化自适应的情感TTS框架,采用交互式遗传算法对个体A-V感知空间进行交互式优化,通过调整每个听者的情感表征,系统生成的语音比使用平均A-V值的模型具有更符合感知的情感表达,对日本、中国、印尼参与者的评估凸显了个性化与文化自适应对超越一刀切情感TTS的重要性。
英文摘要
The rise of conversational AI has increased interest in emotional Text-to-Speech (TTS). Most systems rely on discrete emotion labels, which fail to capture the nuanced nature of human affect. Recent models employ dimensional representations such as Russell's arousal-valence (A-V) model, offering finer control. However, emotional perception varies across individuals and cultures, which may cause mismatches between modeled and perceived emotions. We propose a personalized and culturally adaptive emotional TTS framework that performs interactive optimization of individualized A-V perception spaces using an Interactive Genetic Algorithm. By adapting emotion representations to each listener, the system produces speech with more perceptually aligned emotional expression than models using averaged A-V values. Evaluations with Japanese, Chinese, and Indonesian participants highlight the importance of personalization and cultural adaptation for moving beyond one-size-fits-all emotional TTS.
CommentsAccepted for publication at INTERSPEECH 2026