发表机构
HES-SO; SIB, Swiss Institute of Bioinformatics(瑞士西部应用科学与艺术大学; 瑞士生物信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TransBERT提出仅用合成翻译数据预训练语言模型,在法语生命科学领域达到最先进性能,并发布工具包、语料库和模型。
AI 中文摘要
在专业领域中,非英语语言数据的稀缺严重限制了有效自然语言处理(NLP)工具的开发。我们提出了TransBERT,一个仅使用合成翻译文本进行语言模型预训练的新框架,并介绍了TransCorpus,一个可扩展的翻译工具包。聚焦于法语生命科学领域,我们的方法表明,仅利用合成翻译数据即可在各种下游任务上实现最先进的性能。我们发布了TransCorpus工具包、TransCorpus-bio-fr语料库(36.4GB的法语生命科学文本)、TransBERT-bio-fr及其相关的预训练语言模型,以及用于预训练和微调的可复现代码。我们的结果凸显了在高资源翻译方向上使用合成翻译来构建低资源语言/领域对中高质量NLP资源的可行性。
英文摘要
The scarcity of non-English language data in specialized domains significantly limits the development of effective Natural Language Processing (NLP) tools. We present TransBERT, a novel framework for pre-training language models using exclusively synthetically translated text, and introduce TransCorpus, a scalable translation toolkit. Focusing on the life sciences domain in French, our approach demonstrates that state-of-the-art performance on various downstream tasks can be achieved solely by leveraging synthetically translated data. We release the TransCorpus toolkit, the TransCorpus-bio-fr corpus (36.4GB of French life sciences text), TransBERT-bio-fr, its associated pre-trained language model and reproducible code for both pre-training and fine-tuning. Our results highlight the viability of synthetic translation in a high-resource translation direction for building high-quality NLP resources in low-resource language/domain pairs.
Comments17 pages
Journal refFindings of the Association for Computational Linguistics: EMNLP 2025, pages 19338-19354
DOI:10.18653/v1/2025.findings-emnlp.1053