ReMova:针对英语到白俄罗斯语翻译的大语言模型微调
ReMova: Fine-tuning LLMs for English to Belarusian translation
- University of Oslo(奥斯陆大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出白俄罗斯语专用数据清洗流程,用于微调英语-白俄罗斯语机器翻译模型,实验表明过滤数据对LLM模型收益显著,数据质量是主要瓶颈。
中文摘要 AI 辅助
本文提出了一种针对白俄罗斯语的专用数据清洗流程,并用于英语-白俄罗斯语机器翻译的微调。我们的清洗流程与其他流程的不同之处在于,采用了一个修正工具来处理白俄罗斯语两种正字法的问题、训练数据中的噪声、其他语言的干扰以及白俄罗斯语在互联网上常见的其他拼写错误。在未过滤的训练数据上进行匹配消融实验表明,对所有微调模型而言,过滤带来了显著收益,其中基于大语言模型的模型从过滤中获得的收益大约是专用编码器-解码器机器翻译系统的两倍,这支持了以下观点:对于白俄罗斯语机器翻译,主要瓶颈之一是数据质量。
英文摘要
This paper presents a Belarusian-specific data-cleaning pipeline and fine-tuning for English-Belarusian machine translation. Our cleaning pipeline distinguishes itself from others by employing a correction tool that addresses the issue of the two orthographies of the Belarusian language, noise in the training data, interference from other languages and other misspelling issues common in Belarusian on the internet. A matched ablation on unfiltered training data shows substantial benefits from filtering for all fine-tuned models, with the LLM-based models gaining roughly twice as much from filtering as the dedicated encoder-decoder MT system, supporting the view that for Belarusian MT one of the primary bottlenecks is data quality.