arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38157cs.SDcs.CLeess.AS

EmoRES-TTS:残差增强向量引导的情感语音生成

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra

首次发表
浏览论文内容

中文总结 AI 辅助

针对情感TTS可控性差且训练代价高的问题,提出免训练的EmoRES方法,将情感向量分解为共享与残差分量分别控制,在IEMOCAP上显著提升情感表达准确率与自然度。

中文摘要 AI 辅助

情感条件文本到语音(TTS)模型可能无法可靠地表达所请求的情感,而通过额外训练来提高可控性在计算和情感标注语音训练数据方面都代价高昂。因此,我们研究向量引导(vector steering),这是一种免训练的方法,可修改冻结模型的内部表示。CoCoEmo是一种用于情感TTS的传统向量引导方法,它将每个情感向量视为由单一全局强度控制的不可分割方向,从而限制了其对所请求情感的遵循程度。在这项工作中,我们首先发现情感向量可以分解为一个共享分量(将语音从中性表达移开)和一个残差分量(将生成导向所请求的情感)。基于这一发现,我们提出了情感残差增强引导(Emotion Residual-Enhanced Steering for TTS,EmoRES),这是一种无需重新训练骨干网络即可控制这两个分量的新方法。在IEMOCAP上,EmoRES在IndexTTS-2和CosyVoice2骨干网络上,在所有四个客观情感指标上均优于CoCoEmo。秩相关分别提高了26.13和12.97个百分点,对应相对增益为118.8%和33.1%;情感命中率分别提高了12.95和6.92个百分点,对应相对增益为20.1%和9.8%。人工评估进一步显示,听者正确识别主导请求情感的比例相对提升最高达35.0%,保真度相对提升最高达17.3%,而在成对比较中,听者在高达63.8%的情况下更偏好EmoRES的自然度。组件消融进一步表明,有效的控制得益于保留共享分量,同时增强情感引导向量的残差分量。

英文摘要

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.

发表机构

  • Reality Labs at Meta(Meta现实实验室)
  • National Taiwan University(国立台湾大学)
  • FAIR at Meta(Meta基础人工智能研究团队)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑