AI 中文总结
本文提出一种通过残差映射器干预发音对齐潜在空间的方法,在保留说话人和韵律信息的同时增强构音障碍语音,并验证了显式发音监督优于无监督表征。
AI 中文摘要
构音障碍语音重建(DSR)通常依赖于从受损语音中提取的语言或音素表征来生成更清晰的波形。我们研究了一种互补方法,即干预与语音产生相关的表征。从神经分析-合成框架出发,我们引入了一个残差映射器,它在保留说话人和韵律信息的同时修改发音对齐的潜在空间。该映射器在平行的合成健康语音和人工构音障碍语音上进行预训练,然后使用音素引导和对抗性目标适应到自然构音障碍语音。我们进一步将发音瓶颈与维度匹配的无监督表征进行比较,以评估显式发音监督的益处。
英文摘要
Dysarthric speech reconstruction (DSR) typically relies on linguistic or phonetic representations extracted from impaired speech to generate a more intelligible waveform. We investigate a complementary approach that instead intervenes in a representation related to speech production. Starting from a neural analysis--synthesis framework, we introduce a residual mapper that modifies an articulatory-aligned latent space while preserving speaker and prosodic information. The mapper is pretrained on parallel synthetic healthy and artificially dysarthric speech, then adapted to natural dysarthric speech using phoneme-guided and adversarial objectives. We further compare the articulatory bottleneck with a dimension-matched unsupervised representation to assess the benefit of explicit articulatory supervision.