发表机构
Graduate School of Informatics, Kyoto University(京都大学信息学研究科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出音素引导初始化方法,通过预训练音频编码器于S2P、大语言模型于P2G任务,再端到端微调,在低资源语音识别上匹配或超越级联及端到端基线。
AI 中文摘要
语音大语言模型(speech LLMs)在拥有充足配对语音-文本数据时,在自动语音识别(ASR)任务上表现良好,但在低资源场景下性能会下降。研究表明,先进行语音到音素(S2P)转换,再进行音素到字素(P2G)转换的级联流程,在此类场景下优于端到端语音大语言模型,这表明在配对数据稀缺时,音素中介处理是有益的。我们提出音素引导初始化(phoneme-guided initialization),一种在端到端框架内利用这一见解的简单方法:我们在S2P任务上预训练音频编码器,在P2G任务上预训练大语言模型,然后将两者连接,并在目标ASR任务上对完整模型进行端到端微调。在日语(CSJ)、中文(AISHELL-1)以及Common Voice 25.0中的两种低资源语言(鞑靼语和乌尔都语)上的实验表明,我们的方法在性能上匹配或超越了级联S2P-P2G基线和未使用P2G初始化的端到端模型。
英文摘要
Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.
CommentsAccepted at IEEE SLT 2026