arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26925cs.CLeess.AS

仅使用视觉 grounding 将不同语言的书面词映射到口语词

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

  • Politehnica Bucharest(布加勒斯特理工大学)
  • Stellenbosch University(斯坦陵布什大学)
  • Trinity College Dublin(都柏林圣三一学院)

机构由 AI 辅助整理,请以论文原文为准。

Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper

中文总结 AI 辅助

本研究针对低资源场景下的语音数据构建问题,提出一种基于自监督语音表示的对齐方法,可从视觉 grounding 数据中直接学习跨语言词到语音的映射,效果优于以往的注意力模型。

中文摘要 AI 辅助

在许多低资源场景中,即便是收集语音数据也十分困难,一个有前景的方法是让说话人描述图像,但如何从这类视觉 grounding 的语音数据中构建模型呢?给定带有印地语口语字幕的图像数据集,我们考虑如何将书面英语关键词映射到该词在印地语中的口语实现。以往研究训练端到端多模态神经模型,而我们探索一种基于自监督语音表示的更简单的对齐方法:使用现成的图像字幕系统从图像自动获取书面英语标签,再利用自监督特征对同一关键词对应的印地语话语进行对齐,聚合对齐证据以识别对应目标词的重复语音片段。对关键词 spotting 和定位的实验表明,我们的对齐方法优于以往基于注意力的神经模型,还显示了在对齐过程中纳入负例的益处。本研究证明,无需转录或显式模型训练,即可直接从视觉 grounding 学习跨语言词到语音的映射。

英文摘要

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speech data? Given a dataset of images with Hindi spoken captions, we consider how we can map a written English keyword to spoken realisations of that word in Hindi. Previous work trained end-to-end multimodal neural models. Instead, we explore a simpler alignment-based approach built on self-supervised speech representations. Written English tags are automatically obtained from images using off-the-shelf image captioning systems. Hindi utterances associated with the same keyword are then aligned (using self-supervised features), and alignment evidence is aggregated to identify recurring speech segments corresponding to the target word. Experiments evaluating keyword spotting and localization show that our alignment-based approach outperforms a previous attention-based neural model. We also show the benefit of incorporating negative examples during alignment. Our work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.

补充信息

↑