冻结语音表示中的部分口音控制编辑用于口音转换
Partial Accent-Control Editing in Frozen Speech Representations for Accent Conversion
浏览论文内容
中文总结 AI 辅助
本文提出PACE框架,通过编辑冻结WavLM表示并融合目标口音参考特征,实现可控强度的口音转换,在口音转换与源保留上优于基线。
中文摘要 AI 辅助
口音转换是修改语音录音的任务,使其听起来更接近目标口音,同时保留语言内容和其他与说话者相关的特征。大多数口音转换系统使用经过训练的生成模型。尽管这些模型能够诱导出目标口音语音,但口音修改的强度在推理时不可控,这使得分析口音转换强度如何影响源语音保留变得困难。我们提出了部分口音控制编辑(PACE),一种基于冻结WavLM表示编辑的口音转换框架,无需训练口音条件生成器。首先对源WavLM特征施加受约束的编辑,然后将其与从非平行口音示例中检索到的目标口音参考特征进行融合。融合权重用于控制口音转换强度与其他源属性退化之间的权衡。使用一套包含五个指标的评价体系,并以口音纠缠的说话者相似度为代价,我们展示了PACE在口音转换和源语音保留方面均显著优于一对竞争性基线。
英文摘要
Accent conversion is the task of modifying a speech recording so that it sounds closer to a target accent while preserving linguistic content and other speaker-related characteristics. Most accent conversion systems use trained, generative models. Although they can induce target-accented speech, the strength of accent modification is not controllable at inference time, making it difficult to analyse how the strength of accent conversion affects source preservation. We propose Partial Accent-Control Editing (PACE), an accent conversion framework based upon the editing of frozen WavLM representations without the training of an accent-conditioned generator. A constrained edit is first applied to source WavLM features, which are then fused with target-accent reference features retrieved from non-parallel accent examples. Fusion weights are used to control the trade-off between accent conversion strength and the degradation of other source attributes. Using a suite of five metrics, and with the cost of accent-entangled speaker similarity, we show that PACE is substantially superior to a pair of competitive baselines in terms of both accent conversion and source preservation.
发表机构
- EURECOM Sophia Antipolis, France(欧洲通信研究所)
机构由 AI 辅助整理,请以论文原文为准。