arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEmoEdit:探测与利用预训练语音流的情感可编辑性

SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows

Tianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao, Li Liu

arXiv 2609.34648首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SEmoEdit,首个无需训练的语音情感编辑框架,利用预训练TTS模型的流匹配动力学实现情感替换、擦除与插值,并通过SEmoEditBench验证其高效性与广泛适用性。

AI 中文摘要

现有的基于训练的语音情感编辑方法通常需要大量的任务特定训练,且可能不稳定。这促使我们探究大规模文本到语音(TTS)模型的预训练生成动力学是否可以直接被操控,以实现无需训练的情感编辑。为回答此问题,我们通过构建一个受控测试集,并沿生成轨迹系统诊断编辑效果,来探测预训练的流匹配和混合TTS模型的可编辑性。我们的分析揭示,预训练TTS模型在情感方面具有显著的可编辑性,但这种可编辑性依赖于架构和轨迹,且可能被早期的流匹配步骤破坏,而跨说话人情感迁移会携带除情感外的额外声学属性。为解决这些局限,我们提出SEmoEdit,这是首个无需训练的框架,将情感编辑表述为源情感与目标情感之间的动态速度传输,从而直接在预训练TTS模型内实现稳健的、基于流的情感编辑。SEmoEdit统一了三种核心操作:情感替换、情感擦除和连续情感插值,无需参数更新或任务特定优化。为系统评估这些能力,我们引入了SEmoEditBench,一个包含600个编辑案例的数据集,并在最先进的(SOTA)模型和骨干网络上进行了广泛实验。结果表明,SEmoEdit高效且广泛适用,优于现有的基于训练和激活引导的方法。最终,这项工作揭示了预训练语音流蕴含丰富的情感编辑潜力,为实际应用提供了有益指导。代码、基准和音频样本可在该https URL获取。

英文摘要

Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion editing. To answer this question, we probe the editability of pretrained flow-matching and hybrid TTS models by constructing a controlled test set and systematically diagnosing editing effects along the generative trajectory. Our analysis reveals that pretrained TTS models are substantially editable in emotion, but such editability is architecture- and trajectory-dependent and can be disrupted by early flow-matching steps, while cross-speaker emotion transport carries additional acoustic attributes beyond emotion. To address these limitations, we propose SEmoEdit, the first training-free framework that formulates emotion editing as dynamic velocity transport between source and target emotions, enabling robust, flow-based speech emotion editing directly within pretrained TTS models. SEmoEdit unifies three core operations: emotion replacement, emotion erasure, and continuous emotion interpolation, requiring neither parameter updates nor task-specific optimization. To systematically evaluate these capabilities, we introduce SEmoEditBench, a dataset comprising 600 editing cases, and conduct extensive experiments across state-of-the-art (SOTA) models and backbones. Our results show that SEmoEdit is highly effective and broadly applicable, outperforming existing training-based and activation-steering methods. Ultimately, this work reveals that pretrained speech flows possess rich, latent emotion-editing capabilities, providing useful guidance for real applications. Code, benchmark, and Audio samples are available at https://github.com/imxtx/SEmoEdit.

Comments25 pages, 12 figures, 17 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑