发表机构
MIT CSAIL; MIT RLE; KAIST GSCT(麻省理工学院计算机科学与人工智能实验室; 麻省理工学院电子学研究实验室; 韩国科学技术院文化技术研究生院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种可微且GPU加速的声道声学模拟器,通过梯度下降从语音逆推声道形状,并集成湍流模型与神经网络参数化,实现跨语言自监督自编码及无配对数据的语音驱动MRI重建。
AI 中文摘要
声道是人体中负责过滤声音以产生语音的区域。在本文中,我们提出了一种用于声道的可微且GPU加速的声学模拟器。该可微模拟器通过沿声道的声学管模型传播声音来合成语音,并利用其梯度来解决逆问题:仅从声道产生的声音中重建声道的形状。尽管几何与声音之间的逆映射以非凸性著称,但我们发现梯度下降在以下三项技术贡献下能够成功:(1)我们设计了声道流体动力学的频域公式,其GPU并行化程度比时域有限差分法高70倍;(2)我们集成了湍流的可微模型以合成辅音;(3)与先前在隐式神经表示(INRs)和神经场方面的工作类似,我们发现用神经网络参数化几何形状能加速收敛并避开使离散表示陷入局部极小值的陷阱。由于该模拟器是可微的,它可以轻松集成到其他深度学习流程中,从而支持新颖的语言学和医学成像应用。(1)我们展示了跨11种语言的声道形状的自监督自编码;(2)我们将模拟器与MRI(磁共振成像)图像的生成模型耦合,仅凭语音即可重建一个人的动态声道,而无需配对数据。
英文摘要
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.
CommentsAccepted as NeurIPS 2026 spotlight paper. Supplementary material at https://people.csail.mit.edu/echen/vocal_recon/