arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22300cs.CL

通过跨语言转移和LoRA适配器合并实现低资源阿拉伯语脚本语言的生物医学机器翻译

Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging

Abdullah Alabdullah, Arash Eslamighayour, Sarp Harbalioglu, Lifeng Han

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对阿拉伯语脚本低资源语言生物医学翻译稀缺问题,利用阿拉伯语和波斯语作枢纽,通过LoRA微调训练特定领域适配器,评估少样本学习、最小监督适应及零数据LoRA适配器合并三种策略,展现不同策略效果及跨语言转移局限性。

中文摘要 AI 辅助

我们对医疗领域的跨语言转移进行了系统研究,以解决阿拉伯语脚本语言生物医学神经机器翻译资源稀缺的问题。我们使用阿拉伯语和波斯语作为高资源枢纽,以改善四种严重低资源目标语言的翻译:达里语(阿富汗波斯语,波斯语的一种标准化变体)、普什图语、索拉尼库尔德语(库尔德语的主要标准化变体)和乌尔都语(与印地语密切相关)。我们在仅含解码器的小型语言模型上使用LoRA微调,训练特定领域的枢纽适配器,并评估三种转移策略:少样本上下文学习、最小监督适应,以及据我们所知在此设置下首次使用的零数据LoRA适配器合并。仅用500个句子进行监督适应,达里语(CHrF++ 41.01)就能达到接近枢纽语言的质量,乌尔都语也有显著提升(28.88),而适配器合并在不增加额外成本的情况下,使达里语与监督适应的差距缩小到3.5 CHrF++以内。普什图语和索拉尼库尔德语在高风险临床部署中仍不足,这暴露了与枢纽语言结构距离过大时跨语言转移的局限性。LoRA适配器合并对密切相关的语言效果惊人,即使没有目标语言的生物医学数据。

英文摘要

We present a systematic study of healthcare-domain cross-lingual transfer to address the scarcity of biomedical NMT resources for Arabic-script languages. We use Arabic and Persian as higher-resource pivots to improve translation for \textbf{four severely low-resource} targets: Dari (Afghan Persian, a standardised variety of Persian), Pashto, Sorani Kurdish (Central Kurdish, a major standardized variety of Kurdish), and Urdu (closely related to Hindi). Using LoRA fine-tuning on small decoder-only LLMs, we train \textit{domain-specific pivot adapters} and evaluate \textbf{three transfer strategies}: few-shot in-context learning, minimal supervised adaptation, and, to the best of our knowledge, for the first time in this setting, zero-data LoRA adapter merging. Supervised adaptation with just 500 sentences achieves near pivot-language quality for Dari (CHrF++ 41.01) and meaningful gains for Urdu (28.88), while adapter merging reaches within 3.5 CHrF++ of supervised adaptation for Dari at zero additional cost. Pashto and Sorani Kurdish remain insufficient for high-stakes clinical deployment exposing the limits of cross-lingual transfer when structural distance from the pivots is too great. LoRA adapter merging works surprisingly well for closely related languages, even without target-language biomedical data.

↑