跨方言孟加拉语区域方言命名实体识别:基于留一方言交叉验证与可解释人工智能
Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI
浏览论文内容
中文总结 AI 辅助
针对孟加拉语五种区域方言,提出基于LODOCV策略的跨方言NER框架,Multilingual-E5 Large取得最佳性能(F1最高97.26%),并利用LIME揭示模型决策依赖实体词表面形式。
中文摘要 AI 辅助
孟加拉语是世界第七大语言,具有显著的区域方言多样性,其中巴里萨尔、吉大港、锡尔赫特、诺阿卡利和迈门辛等方言在词汇、形态和句法特征上存在差异。这些差异给命名实体识别(NER)带来了巨大挑战,限制了在标准孟加拉语或单一区域方言上训练的模型的泛化能力。本文提出了一个跨方言孟加拉语NER框架,使用公开可用的ANCHOLIK-NER数据集,该数据集包含五种主要孟加拉语区域方言的17,405条带注释句子和101,817个词元。采用留一方言交叉验证(LODOCV)策略,在四种方言上训练模型,并在剩余的未见方言上进行评估。在相同实验设置下,评估了八种基于预训练Transformer的模型,包括BanglaBERT、MuRIL、XLM-RoBERTa和Multilingual-E5。Multilingual-E5 Large在每一折中均取得最高F1分数,在迈门辛方言上达到峰值97.26%,在最具挑战性的目标方言吉大港方言上达到最低值82.38%。为提高可解释性,将局部可解释模型无关解释(LIME)应用于词级预测,揭示模型的决策主要受目标实体词本身的表面形式驱动,而非周围句子上下文。这些发现为跨方言孟加拉语NER建立了基准,并证明了基于Transformer的迁移学习对低资源区域方言的有效性,同时指出现有模型对表面形式的依赖作为未来工作的方向。
英文摘要
Bangla, the seventh most spoken language in the world, exhibits significant regional dialectal diversity, with dialects such as Barishal, Chattogram, Sylhet, Noakhali, and Mymensingh differing in lexical, morphological, and syntactic characteristics. These variations pose substantial challenges for Named Entity Recognition (NER), limiting the generalization of models trained on Standard Bangla or a single regional dialect. This paper presents a cross-dialect Bangla NER framework using the publicly available ANCHOLIK-NER dataset, comprising 17,405 annotated sentences and 101,817 tokens across five major Bangla regional dialects. A Leave-One-Dialect-Out Cross-Validation (LODOCV) strategy is adopted, training models on four dialects and evaluating on the remaining unseen dialect. Eight pretrained transformer-based models, including BanglaBERT, MuRIL, XLM-RoBERTa, and Multilingual-E5, are evaluated under identical experimental settings. Multilingual-E5 Large achieves the highest F1-score in every fold, peaking at 97.26% on Mymensingh and reaching its lowest, 82.38%, on Chattogram, the most challenging target dialect. To improve interpretability, Local Interpretable Model-agnostic Explanations (LIME) are applied to word-level predictions, revealing that the model's decisions are driven primarily by the surface form of the target entity word itself rather than by surrounding sentence context. These findings establish a benchmark for cross-dialect Bangla NER and demonstrate the effectiveness of transformer-based transfer learning for low-resource regional dialects, while highlighting the surface-form dependence of current models as a direction for future work.
发表机构
- Ahsanullah University of Science and Technology(阿赫萨努拉科技大学)
- Southeast University(东南大学)
- Technische Universität Dresden(德累斯顿工业大学)
- Bangladesh University of Engineering and Technology(孟加拉国工程技术大学)
机构由 AI 辅助整理,请以论文原文为准。