发表机构
Leiden University(莱顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ARIS提出一种结合神经参数估计与确定性DSP合成的神经声源-滤波器模型,实现低资源下对语音声学线索的精确操控,在小型多语种语料库上重合成质量媲美WORLD,参数编辑准确性优于Praat KlattGrid。
AI 中文摘要
语音学家通常需要构建这样的刺激:在保持良好语音质量的同时,精确操控特定的声学线索。经典合成方法与现代神经方法处于精确参数控制与高保真度之间的权衡之中,而神经合成通常需要比语音学家易于获取的更多数据。我们提出了ARIS(解析共振可解释合成),一种神经声源-滤波器模型,它将神经参数估计与确定性DSP合成相结合。每个控制参数都是合成器的一个系数,因此基频(F0)、共振峰和声门声源可以直接编辑。在三种语言的五个小型单说话人语料库上,ARIS重合成的语音质量与WORLD相当,且编辑单个参数的准确性优于Praat KlattGrid,线索之间的串扰可忽略不计。与在大语料库上预训练并在相同数据上微调的HiFi-Glot相比,ARIS在预测自然度上得分略低,但更忠实地再现了录音,并更精确地操控了它们。音频样本:此https URL。
英文摘要
Phoneticians often need to construct stimuli in which specific acoustic cues are precisely manipulated while preserving decent speech quality. Classical synthesis and modern neural methods sit along a trade-off between precise parametric control and high fidelity, and neural synthesis typically demands more data than phoneticians can easily obtain. We present ARIS (Analytic Resonant Interpretable Synthesis), a neural source-filter model that pairs neural parameter estimation with deterministic DSP synthesis. Every control is a coefficient of the synthesizer, so F0, formants and the glottal source can be edited directly. On five small single-speaker corpora in three languages, ARIS resynthesizes speech with quality comparable to WORLD and edits single parameters more accurately than Praat KlattGrid, with negligible crosstalk between cues. Compared with HiFi-Glot, pre-trained on a large corpus and fine-tuned on the same data, ARIS scores slightly lower on predicted naturalness but reproduces the recordings more faithfully and manipulates them more precisely. Audio samples: https://n1r.github.io/ARIS_nsf/.
Comments5 pages, 2 figures, 3 tables. Submitted to ICASSP 2027. Audio samples: https://n1r.github.io/ARIS_nsf/