arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

发音器官源-滤波器文本转语音:通过声道运动学实现物理可控

Articulatory Source-Filter TTS: Physically Grounded Control through Vocal Tract Kinematics

Jesuraj Bandekar, Shinji Watanabe, Prasanta Kumar Ghosh

arXiv 2610.00735首次发表:更新:

发表机构

Indian Institute of Science (IISc); Carnegie Mellon University(印度科学学院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于发音运动学的可控源-滤波器TTS,通过AAI生成轨迹控制滤波器,OT-CFM细化源,实现解耦与细粒度控制如口音修改。

AI 中文摘要

现代神经文本转语音(TTS)系统在声学保真度上表现出色,但作为黑箱,对声道滤波器的可解释控制能力有限。我们提出了一种基于发音运动学的可控源-滤波器TTS架构。一个由大规模预训练表示增强的声学到发音逆推(AAI)模型,为大型TTS语料库生成运动学伪轨迹。这些轨迹调节滤波器响应,而预测的基频和能量轮廓参数化声门源。源和滤波器独立预测,源通过最优传输条件流匹配(OT-CFM)模块细化,然后重组为最终频谱图。我们的模型在可理解性和自然度上与相似规模的基线相当,仅以适度的频谱保真度成本换取显式控制。评估揭示了清晰的源-滤波器解耦,实现稳定的韵律缩放和跨说话人源/滤波器重组,其中F0保持与源说话人绑定,而声道特征跟随滤波器说话人。最后,直接操作发音轨迹可实现细粒度控制,如口音修改,为可解释语音合成提供了新方向。音频样本:此https URL

英文摘要

Modern neural text-to-speech (TTS) systems achieve remarkable acoustic fidelity but act as black boxes, offering little interpretable control over the vocal tract filter. We propose a controllable source-filter TTS architecture grounded in articulatory kinematics. An Acoustic-to-Articulatory Inversion (AAI) model, enhanced by large-scale pretrained representations, generates kinematic pseudo-trajectories for a large TTS corpus. These trajectories condition the filter response, while predicted pitch and energy contours parameterise the glottal source. Source and filter are predicted independently, and the source is refined by an Optimal Transport Conditional Flow Matching (OT-CFM) module before recombination into the final spectrogram. Our model achieves intelligibility and naturalness competitive with similarly sized baselines, with only a modest spectral fidelity cost in exchange for explicit control. Evaluations reveal clear source-filter disentanglement, enabling stable prosodic scaling and cross-speaker source/filter recombination, where F0 remains tied to the source speaker while vocal-tract characteristics follow the filter speaker. Finally, direct manipulation of articulatory trajectories enables fine-grained control, such as accent modification, offering a new direction for interpretable speech synthesis. Audio samples: https://coding-phoenix-12.github.io/ArticulatorySFTTS/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑