arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40041cs.CL

MGhana-ST:面向加纳语言的低资源语音翻译数据集及多语言训练权衡分析

MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs

  • Ashesi University(阿什西大学)
  • AdwumaTech AI
  • HSE University(高等经济大学)
  • GIFT International Fintech Institute(GIFT国际金融科技学院)

机构由 AI 辅助整理,请以论文原文为准。

Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson, Kishankumar Bhimani, Jochebed Afua Basil

中文总结 AI 辅助

本文提出MGhana-ST数据集,涵盖四种加纳低资源语言,发现数据稀缺下多语训练无益甚至有害,并揭示单次运行迁移结果不可靠。

中文摘要 AI 辅助

我们提出了MGhana-ST,一个面向四种低资源加纳语言变体的语音翻译数据集:Ga、Twi(阿库阿佩姆和阿桑特)、Ewe和Fante。MGhana-ST是一项持续进行的标注工作;本文实验使用了约16.1小时的配对语音和英语翻译的固定子集。音频选自两个现有的加纳语音资源。与这些资源不同,英语翻译由37名母语标注者直接根据音频生成,并包含言语和非言语事件标注。使用Whisper-small,我们在严重数据稀缺条件下比较了单语和多语训练,报告了三个随机种子的平均值。在此条件下,平坦多语训练对任何语言变体均无益处。Ga和Twi在种子方差内保持不变(相对于单语标准差1.63和2.20,BLEU分别为+0.51和+0.06),而Ewe下降了6.99 BLEU,Fante下降了5.11。性能下降的语言变体是Ewe(其语言学上独特且来自不同的源语料库)和Fante(资源最少的语言)。将经验性跨语言迁移与基于类型学的相似性进行比较,我们发现迁移BLEU比URIEL相似性更能识别紧密交互的语言对,尽管两者都无法预测哪些语言变体从联合训练中受益。我们还报告了一个方法论发现。早期的单次运行分析发现四种语言变体中有三种存在正迁移;这在跨种子复制中未能成立。对于Ga和Twi,基于1.6至6.2小时音频训练的单语基线,其种子标准差约为多语模型的五倍和三十倍(分别为0.35和0.07 BLEU)。当单语条件噪声更大时,单次运行比较仅凭种子变异即可显示这种规模的表观迁移。我们发布MGhana-ST以支持非洲语言语音技术和低资源语音翻译的研究。

英文摘要

We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The audio is curated from two existing Ghanaian speech resources. Unlike in those resources, the English translations are produced directly from audio by 37 native-speaker annotators and include verbal and non-verbal event annotations. Using Whisper-small, we compare monolingual and multilingual training under severe data scarcity, reporting means over three seeds. Flat multilingual training benefits no variety in this regime. Ga and Twi are unchanged within seed variance (+0.51 and +0.06 BLEU against monolingual standard deviations of 1.63 and 2.20), while Ewe declines by 6.99 BLEU and Fante by 5.11. The degrading varieties are Ewe, which is linguistically distinct and drawn from a different source corpus, and Fante, the least-resourced. Comparing empirical cross-lingual transfer with typology-based similarity, we find that transfer BLEU identifies closely interacting language pairs better than URIEL similarity, though neither predicts which varieties benefit from joint training. We also report a methodological finding. An earlier single-run analysis found positive transfer for three of four varieties; this did not survive replication across seeds. For Ga and Twi, monolingual baselines trained on 1.6 to 6.2 hours of audio have seed standard deviations roughly five and thirty times those of the multilingual models (0.35 and 0.07 BLEU). When the monolingual condition is noisier, a single-run comparison can show apparent transfer of this size from seed variation alone. We release MGhana-ST to support research on African language speech technology and low-resource speech translation.

↑