HomoEnsNER:古吉拉特语命名实体识别中语言对齐是否优于架构复杂度?
HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?
浏览论文内容
中文总结 AI 辅助
该研究针对古吉拉特语NER,提出同构集成模型HomoEnsNER,经实验证实其F1值优于基线及异构模型,表明语言对齐比架构复杂度更适合低资源印度语言NER。
中文摘要 AI 辅助
古吉拉特语的命名实体识别(NER)研究仍较为欠缺,其面临的阻碍包括缺乏大小写线索、丰富的形态变化、词汇歧义以及自由词序。以往的集成工作强调通过组合异构分类器、多语言编码器或经典序列模型来实现架构多样性,而非利用语言对齐的单语预训练。本研究针对古吉拉特语这类低资源、形态丰富的语言,探究由单个单语编码器构成的同构集成是否优于此类架构多样性。我们提出HomoEnsNER,这是一个由5个独立微调的GujaratiBERT模型组成的同构集成,通过多数投票进行组合;我们将其与单个GujaratiBERT基线模型以及6种异构替代方案进行评估,这些方案包括与MuRIL-base、MuRIL-large、IndicBERT、mBERT、BiLSTM、CRF的组合,以及堆叠式BiLSTM-CRF-GujaratiBERT架构。所有8个模型均在一致的预算下进行训练,并在Naamapadam古吉拉特语测试集上以实体级F1值进行评估。HomoEnsNER取得了最高的F1值(0.8442),超过了基线模型(0.8347)以及所有异构替代方案(最低为0.7855),这表明对于低资源印度语言的NER而言,语言对齐是一种比架构复杂度更有效、更具成本效益的集成策略。
英文摘要
Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, and free word order. Prior ensemble work has emphasized architectural diversity by combining heterogeneous classifiers, multilingual encoders, or classical sequence models, rather than exploiting language-aligned monolingual pretraining. This study asks whether, for a low-resource, morphologically rich language like Gujarati, a homogeneous ensemble of a single monolingual encoder outperforms such architectural diversity. We propose HomoEnsNER, a homogeneous ensemble of five independently fine-tuned GujaratiBERT models combined via majority voting, evaluated against a single GujaratiBERT baseline and six heterogeneous alternatives, including combinations with MuRIL-base, MuRIL-large, IndicBERT, mBERT, BiLSTM, CRF, and a stacked BiLSTM-CRF-GujaratiBERT architecture. All eight models were trained under a consistent budget and evaluated using entity-level F1 on the Naamapadam Gujarati test split. HomoEnsNER achieved the highest F1 (0.8442), surpassing the baseline (0.8347) and every heterogeneous alternative (lowest: 0.7855), indicating that language alignment is a more effective, budget-conscious ensembling strategy than architectural complexity for low-resource Indian language NER.