arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12018cs.CL

孟加拉语区域方言的多方言神经机器翻译系统

Unified Multi-Dialectal Neural Machine Translation for Bangla Using the Dwadash Benchmark Corpus

Rakib Ullah, Md. Ruhul Islam, Tanbir Ahmed, Nayan Kumar Nath

首次发表
浏览论文内容

中文总结 AI 辅助

针对孟加拉语方言翻译性能差的问题,本研究提出Poly-Dialectal神经机器翻译系统,构建含51531句对的最大多方言语料库,微调BanglaT5模型性能最优,还部署量化网络应用推动数字包容。

中文摘要 AI 辅助

区域方言变异是孟加拉语自然语言处理(NLP)面临的根本挑战,超过2.4亿使用者的各类区域变体,在语音、形态和词汇层面与标准口语孟加拉语(SCB)存在显著差异。当前的神经机器翻译(NMT)架构和大语言模型(LLM)大多假设语言分布同质,导致低资源区域方言翻译时性能严重下降。本研究提出一种统一的Poly-Dialectal神经机器翻译系统,可实现12种孟加拉语区域方言间的多向翻译,无需通过中间标准枢纽。我们构建了迄今最大的孟加拉语多方言平行语料库,包含12种方言的51531条非空平行句对,其中纳入2500条专家验证的双向平行句对,用于5种此前未被研究的方言。在权重分解低秩适配(DoRA)下评估序列到序列架构,我们微调后的BanglaT5模型达到了最先进的翻译性能(29.26 BLEU值、57.26 chrF++值),在保持形态连贯性的同时,优于NLLB-200(6.15亿参数)和mBART-50(6.11亿参数)。此外,我们开展了系统的跨方言迁移分析和语料库规模研究,确立了低资源方言适配的经验阈值。最后,我们将优化后的INT8量化模型部署为开放获取的网络应用,以促进边缘方言群体的数字包容。完整数据集公开于Mendeley Data(https URL)。

英文摘要

Neural Machine Translation (NMT) and Large Language Models (LLMs) excel at cross-lingual tasks but often fail to capture intra-lingual morphological variation, marginalizing dialectal speakers. In Bangla, existing translation frameworks commonly rely on Standard Colloquial Bangla (SCB) as an intermediate pivot, which can compound errors and reduce cross-dialectal nuance. To address this gap, we introduce a unified, multi-directional NMT system for direct translation between SCB and eleven regional variants. We first review prior dialectal NLP resources to identify existing technological gaps. As a foundational contribution, we construct and release a large multi-dialect parallel corpus for Bangla, comprising 14,562 aligned rows and 51,541 non-null sentence pairs through the integration of seven prior datasets and native-speaker-verified manual augmentation. Using this corpus, we benchmark state-of-the-art sequence-to-sequence architectures with parameter-efficient Weight-Decomposed Low-Rank Adaptation (DoRA). Results show that deep monolingual pre-training is more effective than large multilingual capacity for this task. The compact BanglaT5 model outperforms NLLB-200 and mBART-50 by up to 13.96 BLEU, achieving 29.26 BLEU, 57.26 chrF++, and 49.68 METEOR. A dataset scaling study shows diminishing returns beyond 3,000 parallel pairs and indicates that linguistic proximity to Standard Bangla is more important than raw data volume for translation quality. Finally, we deploy the optimized model as an INT8-quantized web application, providing a scalable, open-source framework for inclusive language technology and equitable digital access.

发表机构

  • SYLHET ENGINEERING COLLEGE(锡尔赫特工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑