Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding
仅使用视觉 grounding 将不同语言的书面词映射到口语词
机构 * Politehnica Bucharest(布加勒斯特理工大学) ; Stellenbosch University(斯坦陵布什大学) ; Trinity College Dublin(都柏林圣三一学院)
专题命中 视觉定位与Grounding :grounding(title,title_cn)
AI总结 本研究针对低资源场景下的语音数据构建问题,提出一种基于自监督语音表示的对齐方法,可从视觉 grounding 数据中直接学习跨语言词到语音的映射,效果优于以往的注意力模型。
Comments 9 pages, 5 figures, 5 tables, preprint, submitted to IEEE Transactions on Audio, Speech and Language Processing