EmoSay:用于扩展现实中情感交流的人工智能驱动文本到情感语音系统
EmoSay: Artificial Intelligence-Driven Text-to-Emotional-Speech System for Affective Communication in Extended Reality
AI总结:
该研究提出EmoSay系统,通过离散情感提示调节神经合成流水线,经用户研究验证可显著提升XR沉浸体验与可用性,为包容性XR设计提供了可扩展的感知情感框架。
AI中文摘要:
尽管当前神经文本到语音(TTS)系统已达到很高的可懂度,但它们往往缺乏真实情感交流所需的情感细微差别。这一局限在扩展现实(XR)中尤为关键,因为缺乏情感表现力的音频会降低用户的临场感和空间沉浸感。我们提出了EmoSay,这是一种人工智能驱动的文本到情感语音(TTES)系统,旨在弥合沉浸式环境中的语义-情感鸿沟。EmoSay通过离散情感提示调节神经合成流水线,输出通过基于Unity的界面呈现,该界面具有高保真空间音频。我们通过一项综合用户研究对该系统进行评估,研究重点为感知、参与度和主观同理心。结果表明,EmoSay显著提升了沉浸体验,系统可用性量表(SUS)得分为74.76,表明其可用性强,可无缝集成到XR工作流程中。主观评估显示出高度的感知自然度,且情感表现力与用户参与度之间存在强正相关。回归分析确定,在测试的用户满意度预测因素中,语音自然度是最强的,这表明EmoSay的情感韵律有助于满足沉浸式环境中对真实感的更高期望。这项工作为包容性XR设计贡献了一个可扩展的、感知情感的框架,并证明了合成情感在通过语音优先交互促进人机融洽关系方面的作用。
英文摘要:
While contemporary neural text-to-speech (TTS) systems have achieved high levels of intelligibility, they frequently lack the emotional nuance required for authentic affective communication. This limitation is particularly critical in Extended Reality (XR), where the absence of emotionally expressive audio can diminish user presence and spatial immersion. We present EmoSay, an Artificial Intelligence-driven Text-to-Emotional-Speech (TTES) system designed to bridge the semantic-affective gap in immersive environments. EmoSay modulates a neural synthesis pipeline using discrete emotional prompts, delivering the output through a Unity-based interface featuring high-fidelity spatialized audio. The system was evaluated through a comprehensive user study focusing on perception, engagement, and the subjective sense of empathy. Our results demonstrate that EmoSay significantly enhances the immersive experience, achieving a System Usability Scale (SUS) score of 74.76, indicating strong usability and seamless integration within the XR workflow. Subjective assessments reveal a high degree of perceived naturalness and a strong positive correlation between emotional expressiveness and user engagement. Regression analysis identifies vocal naturalness as the strongest of the tested predictors of user satisfaction, suggesting that EmoSay's affective prosody helps meet the heightened expectations for realism in immersive settings. This work contributes a scalable, affect-aware framework for inclusive XR design and demonstrates the role synthetic emotion can play in fostering human-computer rapport through voice-first interaction.