arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProPS:面向自然语言条件化说话人嵌入分布的提示式画像合成

ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions

Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez, Jesus Villalba Lopez, Najim Dehak

arXiv 2607.05276首次发表:更新:

发表机构

Department of Electrical and Computer Engineering, Johns Hopkins University(约翰霍普金斯大学电气与计算机工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有说话人嵌入提取器仅具备描述性不具备生成性的问题,提出ProPS框架,基于自然语言提示生成符合指定属性的说话人嵌入分布,支撑可控语音生成任务。

AI 中文摘要

说话人嵌入(即x-向量)被广泛用于表征说话人身份及相关属性,但现有嵌入提取器通常仅具备描述性而非生成性:它们将观测到的语音片段映射为x-向量,再将其用于下游任务。本文提出提示式画像合成框架ProPS,可基于“一名带有印度口音的三十岁男性说话人”这类自然语言提示生成说话人嵌入分布。ProPS将人工撰写的画像描述转换为句子嵌入,使用在大规模数据集上训练的混合密度网络预测x-向量空间内的高斯混合模型。该模型通过最大化真实说话人嵌入与目标画像匹配的似然度完成训练,生成的分布通过预留x-向量上的负对数似然、合成采样x-向量的属性分类准确率完成评估。实验表明,ProPS可生成画像条件化分布,输出的x-向量可保留指定的年龄、性别、口音、韵律特征等说话人属性。该设计可为文本转语音(TTS)、语音转换(VC)等语音生成系统提供可控的说话人画像合成能力,同时将生成分布锚定在观测得到的说话人嵌入结构中。

英文摘要

Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications. We introduce ProPS, Prompted Profile Synthesis, a framework for generating distributions of speaker embeddings conditioned on natural language prompts such as "a thirties male speaker with an Indian accent". ProPS converts human-written profile descriptions into sentence embeddings and uses a mixture density network trained on a large-scale dataset to predict a Gaussian mixture model in the x-vector space. The model is trained by maximizing the likelihood that real speaker embeddings match the requested profile, and its generated distributions are evaluated by negative log-likelihood on held-out x-vectors and by attribute classification accuracies on sampled synthetic x-vectors. Experiments show that ProPS produces profile-conditioned distributions and generates x-vectors that preserve requested speaker attributes such as age, gender, accent, and prosodic characteristics. This design enables controllable speaker-profile synthesis for speech generation systems like Text-To-Speech (TTS) or Voice Conversion (VC) while anchoring generated distributions in observed speaker-embedding structure.

CommentsPublished in SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑