arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低资源环境下基于HGNN的跨模态知识迁移增强语音表示学习:以Yemba语为例

Enhancing speech representation learning with cross-modal knowledge transfer with HGNN under low resource settings: the case study of Yemba

Yannick Yomie Nzeuhang, Paulin Melatagia Yonta, Marie Tahon

arXiv 2609.23194首次发表:更新:

发表机构

University of Yaounde I; Le Mans Université; LIUM; IRD; UMMISCO(雅温得第一大学; 勒芒大学; 勒芒大学计算机科学实验室; 法国发展研究所; UMMISCO)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低资源语言声学表示学习数据稀缺问题,提出基于异构图神经网络(HGNN)的跨模态知识迁移方法,通过消息传递实现语言知识向声学节点迁移,实验证明其有效提升声学表示质量。

AI 中文摘要

声学表示学习对于语音处理至关重要,然而低资源语言面临严重的数据稀缺问题,限制了传统方法和自监督方法的有效性。作为一种有前景的替代方案,本工作提出通过基于异构图神经网络的跨模态知识迁移方法来增强声学表示,其中声学实体和语言实体在统一图中被建模为不同的节点类型。通过消息传递机制,语言节点显式地将知识迁移到声学节点,实现结构化且可解释的跨模态信息流。为突出这种知识迁移及其益处,我们测量了标准聚类指标作为声学表示的内在评估,并为了强调适用性,在低资源设置下使用英语基准和喀麦隆语言数据集执行了孤立词识别任务。结果表明,声学表示持续受益于通过图传播的语言知识。据我们所知,这是首次使用HGNN实现声学表示学习的显式跨模态知识迁移,为低资源环境下的语音表示指明了一个有前景的方向。

英文摘要

Acoustic representation learning is crucial for speech processing, yet low-resource languages (LRLs) face severe data scarcity, limiting the effectiveness of traditional and self-supervised methods. As a promising alternative, in this work, we propose to enhance acoustic representation trough a cross-modal transfer knowledge approach, based on heterogeneous graph neural networks (HGNNs), where acoustic and linguistic entities are modeled as distinct node types within a unified graph. Through message-passing mechanisms, linguistic nodes explicitly transfer knowledge to acoustic nodes, enabling structured and interpretable cross-modal information flow. To highlight this knowledge transfer and its benefits, we measured standard clustering metrics as an intrinsic evaluation of acoustic representation, and to emphasize applicability, we performed isolated-word recognition tasks using an English benchmark and a Cameroonian language dataset in low resources settings . Results demonstrate that acoustic representations consistently benefit from linguistic knowledge propagated through the graph. To our knowledge, this is the first demonstration of explicit cross-modal knowledge transfer for acoustic representation learning using HGNNs, highlighting a promising direction for speech representation in low-resource settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑