灵活且可解释的口音距离测量
Flexible and Interpretable Accent Distance Measurements
- University of Cambridge(剑桥大学)
- Cambridge University Press & Assessment(剑桥大学出版社与考评部)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出利用发音反演产生的发音表征和最优传输框架,实现灵活且可解释的口音距离测量,兼顾可解释性与适用性。
AI中文摘要:
确定两位说话者口音之间的差异是语言学和语音技术研究中的一项基本任务。用于测量这些差异的方法取决于具体的研究领域。语音学研究者可能通过比较成对单词录音中的元音共振峰来展示口音变化。这些结果具有可解释性,但录音的收集耗时,且可能无法代表连续语音。带口音的文语转换(TTS)研究已倾向于使用从口音分类任务中提取的口音嵌入。这些嵌入可从任何语音录音中生成,但不易解释。在本文中,我们证明通过发音反演创建的发音表征可用作口音比较的可解释基础,并且最优传输为跨任意录音类型的口音比较提供了框架。
英文摘要:
Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation by comparing vowel formants in paired recordings of individual words. These results will be interpretable, but the recordings will be time-consuming to collect and may not be representative of connected speech. Accented Text-to-Speech (TTS) research has pushed towards using accent embeddings derived from accent classification tasks. These embeddings can be produced from any speech recording, but are not readily interpretable. In this paper, we demonstrate that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.