发表机构
BRAC University; Penta Global Limited; University of Central Florida; Independent University, Bangladesh(BRAC大学; Penta Global有限公司; 中佛罗里达大学; 孟加拉独立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对孟加拉语方言资源匮乏问题,构建首个多标注方言基准5-Dialects-BN,含6,000条五种方言数据,支持方言识别、归一化、翻译等任务,为方言感知NLP提供评估基础。
AI 中文摘要
大语言模型(LLMs)在自然语言处理(NLP)任务中取得了显著进展,但其性能在低资源语言和方言多样性环境中急剧下降。孟加拉语是世界第六大语言,体现了这一差距:现有资源绝大多数针对标准孟加拉语,使其区域方言缺乏开发或评估方言感知系统所需的基准。我们通过5-Dialects-BN弥补了这一空白,这是首个将罗马化音译与方言文本、标准孟加拉语、英语和主观性标签对齐的孟加拉语多标注方言基准,覆盖五种区域变体。该数据集包含6,000条人工标注条目,涵盖五种主要方言:吉大港(Chittagong)、巴里萨尔(Barisal)、诺阿卡利(Noakhali)、锡尔赫特(Sylhet)和朗布尔(Rangpur)(吉大港1,900条;诺阿卡利1,500条;锡尔赫特1,200条;巴里萨尔700条;朗布尔700条),反映了自然的在线可用性。每条条目都配有五种对齐标注:原始方言文本、罗马化音译、英语翻译、标准孟加拉语翻译以及主观性标签(主观vs.客观)。标注由母语者和语言学本科生制作并交叉验证,以确保方言真实性和语义保真度。该资源支持多种任务,包括方言识别、方言到标准归一化、机器翻译、主观性分类以及多语言LLMs的参数高效微调(如LoRA)。通过提供标准化的多标注基准,5-Dialects-BN能够对方言多样的孟加拉语进行原则性LLM评估,并为低资源、方言感知NLP的进一步研究奠定基础。
英文摘要
Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources overwhelmingly target Standard Bangla, leaving its regional dialects without the benchmarks needed to develop or evaluate dialect-aware systems. We address this gap with 5-Dialects-BN, the first multi-annotation Bangla dialect benchmark to align Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five regional varieties. The dataset comprises 6,000 manually annotated entries spanning five major dialects: Chittagong, Barisal, Noakhali, Sylhet, and Rangpur (Chittagong 1,900; Noakhali 1,500; Sylhet 1,200; Barisal 700; Rangpur 700), reflecting natural online availability. Each entry is enriched with five aligned annotations: the original dialectal text, a Romanized transliteration, an English translation, a Standard Bangla translation, and a subjectivity label (subjective vs. objective). Annotations were produced and cross-validated by native speakers and undergraduate linguistics students to ensure dialectal authenticity and semantic fidelity. The resulting resource supports a diverse suite of tasks, including dialect identification, dialect-to-standard normalization, machine translation, subjectivity classification, and parameter-efficient fine-tuning (e.g., LoRA) of multilingual LLMs. By providing a standardized, multi-annotation benchmark, 5-Dialects-BN enables principled evaluation of LLMs on dialectally diverse Bangla and lays a foundation for further research in low-resource, dialect-aware NLP.
Comments31 pages, 18 figures, 26 tables. Accepted to EMNLP 2026 (Main Conference)