arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18739cs.CLcs.AI

低资源语言中自动化NER标注修正的可扩展框架

A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

发表机构芬兰国家技术研究中心 · 穆罕默德·本·扎耶德人工智能大学
查看机构详情
  • VTT Technical Research Centre of Finland Ltd.(芬兰国家技术研究中心)
  • Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

Toqeer Ehsan, Thamar Solorio

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一个多步骤自动化框架,通过频率迭代、自训练和双阈值机制修正低资源语言NER数据集的噪声标注,显著提升性能,并探索了生成式大语言模型的应用潜力。

中文摘要 AI 辅助

在命名实体识别(NER)中,与其他任何NLP任务一样,质量差或带有噪声的标注使得实现最先进的性能变得具有挑战性。在本文中,我们提出了一个多步骤框架,通过采用自动化技术来提高NER数据集的标注质量。我们提出了一种基于频率的迭代方法,利用自训练和双阈值机制来增强推理置信度。在不同NER数据集上的实验评估表明,与原始数据集相比,NER性能有显著提升。这项工作进一步探索了生成式大语言模型(LLMs)在低资源语言中执行NER的潜力。

英文摘要

Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper, we present a multi-step framework to enhance the annotation quality of NER datasets by employing automated techniques. We propose a frequency-based iterative approach that leverages self-training and a dual-threshold mechanism to enhance inference confidence. Experimental evaluations on different NER datasets demonstrate significant improvements in NER performance with respect to the original datasets. This work further explores the potential of generative Large Language Models (LLMs) to perform NER for low-resource languages.

补充信息

↑