转录、翻译与优化:语音翻译的联合奖励学习
Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation
浏览论文内容
中文总结 AI 辅助
针对语音翻译中思维链训练与推理不匹配问题,提出基于GRPO的联合识别与翻译微调方法,在CoVoST 2和FLEURS上显著提升BLEU并降低WER。
中文摘要 AI 辅助
在基于大语言模型的语音翻译中,基于转录的思维链(CoT)存在监督微调(SFT)中使用的参考转录与推理时模型生成的转录之间的不匹配问题。为解决这一问题,我们提出了通过组相对策略优化(GRPO)进行联合识别与翻译微调的方法。我们对转录和翻译均进行评分,其中翻译以模型生成的转录为条件,并比较了三种令牌优势策略。使用Qwen2.5-Omni-3B模型,在四种语言上,我们评估了在SFT和GRPO下CoT与直接语音翻译(Direct ST)的性能,在CoVoST 2上训练,并在CoVoST 2和FLEURS上测试。CoT GRPO在CoVoST 2和FLEURS上的平均BLEU分数分别比Direct ST GRPO高出1.77和0.83分。与CoT SFT相比,GRPO将BLEU提高了0.82和0.67分,并将词错误率(WER)相对降低了8.8%和7.2%。这些结果突显了强化微调作为缓解训练-推理不匹配的有效方法,能够联合提升识别和翻译性能。
英文摘要
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.
发表机构
- University of Iowa(爱荷华大学)
机构由 AI 辅助整理,请以论文原文为准。