解剖感知的完整声道声学-发音反演的跨说话人适配
Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion
浏览论文内容
中文总结 AI 辅助
提出基于解剖标志的几何适配框架,通过仿射变换和TPS变形实现跨说话人声道反演,无需重训,在8位说话人上达到3.19毫米最低误差。
中文摘要 AI 辅助
跨说话人的声学-发音反演需要考虑说话人之间的解剖差异。我们提出了一种几何适配框架,利用解剖标志(主要位于椎骨和牙齿结构上)将固定反演模型的预测迁移到未见过的说话人。仿射变换后接薄板样条(TPS)变形,将10个声道结构的预测轮廓映射到每个目标说话人的几何结构中,无需重新训练。每个说话人选择一个/u/帧作为共同语音参考来识别标志,而不假设跨说话人具有相同的发音配置,所得映射在录音中重复使用。我们在单说话人rt-MRI数据库上训练模型,并在来自独立多说话人rt-MRI数据库的八位说话人上评估适配。我们比较了使用12或14个标志的仿射和TPS配置。Affine12+TPS14实现了最低的平均点到最近点误差3.19毫米。这些结果支持解剖标志信息与非刚性对齐的联合价值。
英文摘要
Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen speakers. An affine transformation followed by thin-plate spline (TPS) deformation maps the predicted contours of 10 vocal-tract structures into each target speaker's geometry without retraining. Landmarks are identified in one selected /u/ frame per speaker as a common phonetic reference without assuming identical articulatory configurations across speakers, and the resulting mapping is reused across recordings. We train the model on a single-speaker rt-MRI database and evaluate adaptation on eight speakers from a separate multi-speaker rt-MRI database. We compare affine and TPS configurations using 12 or 14 landmarks. Affine12+TPS14 achieves the lowest mean point-to-closest-point error of 3.19mm. These results support the combined value of anatomical landmark information and nonrigid alignment.
发表机构
- Université de Lorraine(洛林大学)
- CNRS(法国国家科学研究中心)
- Inria(法国国家信息与自动化研究所)
- Inserm(法国国家健康与医学研究院)
机构由 AI 辅助整理,请以论文原文为准。