arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式模型与编码器模型在多语言命名实体识别(NER)中的对比:针对Naamapadam的全面实证研究

Generative vs. Encoder Models for Multilingual NER: A Comprehensive Empirical Study on Naamapadam

Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar

arXiv 2608.29959首次发表:更新:

发表机构

Rajiv Gandhi University of Knowledge Technologies(拉吉夫·甘地知识技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对Naamapadam基准的11种印度语言,对比生成式与编码器NER模型,发现编码器模型在多数语言上显著优于生成式模型,识别出三类语言集群并给出部署指南。

AI 中文摘要

语言是人类最重要的技术,但印度22种宪法认可语言的超10亿使用者,其数字层面仍存在结构性缺失。命名实体识别(NER)是将原始文本转化为机器可解读知识的基础步骤,针对英语的NER已被深入研究,但多数印度语言的NER问题仍未解决。本文针对Naamapadam基准的全部11种语言,开展生成式与编码器类神经架构的严谨对比研究。我们评估5个经典模型族,涵盖序列到序列Transformer及多语言编码器;4个经LoRA和4位NF4量化微调的仅解码器大语言模型(LLM);以及9个零样本到5样本推理的生成式模型。在严格的CoNLL跨度级评估下,编码器模型(mBERT和XLM-R,印地语F1值均为0.675)在11种语言中的10种,显著优于所有生成式架构,与最强竞争对手Gemma-2-2B的F1平均值0.427相比,差距达7.5至40个百分点。最佳少样本结果仅达到编码器基线的28%。我们识别出三个语言集群:编码器主导型、部分覆盖型和失败区域型,并基于迁移学习和低资源NER原则提供可操作的部署指南。

英文摘要

Language is humanity's most consequential technology, yet for over a billion speakers across India's twenty-two constitutionally recognised languages, its digital layer remains structurally incomplete. Named Entity Recognition (NER), the foundational step in transforming raw text into machine-interpretable knowledge, has been studied exhaustively for English but remains largely unsolved across most Indic languages. This paper presents a rigorous comparative study of generative and encoder-based neural architectures for NER on all eleven languages of the Naamapadam benchmark. We evaluate five classic model families spanning sequence-to-sequence transformers and multilingual encoders; four decoder-only large language models (LLMs) fine-tuned with LoRA and 4-bit NF4 quantisation; and nine generative models in zero-to-5-shot inference. Under strict CoNLL span-level evaluation, encoder-based models (mBERT and XLM-R, both F1=0.675 on Hindi) substantially outperform every generative architecture in ten of eleven languages, with gaps of 7.5-40 percentage points against the strongest competitor (Gemma-2-2B: avg F1=0.427). The best few-shot result reaches only 28% of the encoder baseline. We identify three language clusters--encoder-dominant, partial-coverage, and failure-zone; and provide actionable deployment guidelines grounded in transfer learning and low-resource NLP principles.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑