arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09321cs.CLeess.AS

方言鲁棒的语音语言模型与合成伪方言增强

Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

Shunsuke Mitsumori, Tomoya Mizumoto, Yusuke Fujita

首次发表
浏览论文内容

中文总结 AI 辅助

针对方言数据稀缺导致SLM性能下降的问题,提出用标准TTS合成伪方言语音并引入中间标准文本预测,在日德汉方言翻译上显著提升性能,且无需方言语音资源。

中文摘要 AI 辅助

语音语言模型(SLM)的性能常因数据稀缺而在方言上有所下降。传统的文本到语音(TTS)增强方法难以覆盖多样的方言,因为它需要一定量的真实方言语音。我们提出通过标准语言TTS模型转换由LLM生成的方言文本,从而合成伪方言语音,无需任何真实方言语音。此外,我们在训练过程中引入中间标准文本预测,作为下游任务的语义归一化。我们通过日语、德语和汉语方言的方言到英语语音翻译来评估方言理解能力。与合成标准语音基线相比,伪方言增强提高了日语(从25.38到26.24)和德语(从31.57到32.47)的得分。此外,中间标准文本预测有效弥合了语义差距,将日语性能提升至28.26,汉语性能从11.67提升至16.37。这些结果表明,我们的方法可扩展到各种语言,无需针对每种方言的特定语音资源。

英文摘要

Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.

发表机构

  • SB Intuitions
  • Waseda University(早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑