arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23191cs.CL

低资源跨模态对齐:利用HGNN增强语音表示

Low resource cross-modal alignment using HGNN to enhance speech representation

Yannick Yomie Nzeuhang, Marie Tahon, Paulin Melatagia Yonta

首次发表
浏览论文内容

中文总结 AI 辅助

针对低资源语音-文本对齐,提出基于异构图神经网络和链接预测的数据高效方法,在TIMIT和Yemba语词级任务上达到或超越SAMU-XLSR,且资源消耗更少。

中文摘要 AI 辅助

语音-文本空间对齐是一种多模态表示学习方法,旨在将不同的语音和文本映射到一个共享的表示空间中,从而丰富每种模态的表示。已有的架构,如SAMU-XLSR,通常采用学生/教师框架,目标是对音频编码器进行微调,使其生成的表示与文本表示紧密匹配。通过这种方式,语音表示在语义上得到丰富。然而,此类系统通常需要大量的训练数据和可观的计算资源,这使得它们在资源受限的条件下难以应用于低资源语言。本研究提出了一种基于异构图神经网络和链接预测的数据高效空间对齐方法。其核心思想是利用消息传递机制,将信息从文本模态显式地传递到语音模态,从而减少对大型训练数据集的需求,并以更具可解释性的方式从本质上丰富声学表示。尽管针对高资源语言已进行了深入研究,但对于某些低资源语言,词级语音任务仍然具有重要意义。因此,我们使用TIMIT(英语)数据集和Yemba语(喀麦隆的一种语言)进行了词级语音-文本对齐实验。我们的方法在词检索任务上取得了与最先进方法SAMU-XLSR相当的结果,甚至在Yemba语上超越了它,同时使用的资源要少得多,这证明了其强大性、节俭性和高效性。

英文摘要

Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation space, leading to enrichment of the representation of each modality. Proposed architectures, such as SAMU-XLSR, typically follow a student/teacher framework, with the goal of fine-tuning an audio encoder to produce representations that closely match those of the text. In this way a speech representation is semantically enriched. However, such systems generally require large amounts of training data and considerable computational resource, making them difficult to apply to low resources languages under frugal constraints. The present work proposes a data-efficient space alignment method based on Heterogeneous Graph Neural Networks and link prediction. The core idea is to leverage message passing to explicitly transfer information from the text modality to the speech modality, thereby reducing the need for large training datasets and intrinsically enriching the acoustic representations, all in a more interpretable manner. Although thoroughly explored for high-resource languages, word-level tasks in speech remain relevant for certain low-resource languages. Therefore, we conducted experiments on speech-text alignment at the word level using the TIMIT (English) dataset and Yemba (a Cameroonian language). Our approach yields results comparable to those of SAMU-XLSR, a state-of-the-art method, and even surpasses it for the Yemba language in the task of word retrieval, while using far fewer resources, demonstrating its power, frugality, and efficiency.

发表机构

  • University of Yaounde I(雅温得第一大学)
  • IRD(法国发展研究所)
  • UMMISCO
  • Le Mans Université(勒芒大学)
  • LIUM(勒芒大学计算机科学实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑