arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向含噪大语言模型标签的鲁棒命名实体识别的错误类型感知损失重加权

Error-Type-Aware Loss Reweighting for Robust Named Entity Recognition with Noisy LLM Labels

Elena Merdjanovska, Jonas Golde, Alan Akbik

arXiv 2608.30827首次发表:更新:

发表机构

Humboldt-Universität zu Berlin; Science of Intelligence(柏林洪堡大学; 智慧科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM标注NER数据集存在的异质噪声问题,提出错误类型感知损失重加权方法,无需额外资源即可提升模型F1值,在特定噪声水平下效果显著。

AI 中文摘要

大语言模型正越来越多地用于标注数据集,以训练更小的、任务专门化的模型,例如命名实体识别(NER)。尽管该方法能生成有效的模型,但它假设合成数据集的标注是正确的。在本研究中,我们发现:(i)当前的微调过程完全忽略了大语言模型(LLM)引入的标注噪声,导致性能下降;(ii)现有的抗噪声损失无法迁移到序列标注任务,因为NER中的标注噪声是异质的:例如,提及缺失和类型错误会以不同方式影响训练信号。在抗噪声损失中对所有含噪词元一视同仁,且对所有词元应用单一重加权准则,可能会移除有用的监督信号或强化错误标签。为解决这一局限,我们提出面向NER的错误类型感知损失重加权方法,该方法为不同类型的潜在错误词元引入独立的重加权规则。我们的方法简单高效,无需额外训练资源,在15%至40%的噪声水平下,数据集平均F1值提升0.8至2.0个百分点;在Wikigold数据集24.1%的噪声水平下,最大提升达4.6个百分点。

英文摘要

Large language models are increasingly used to annotate datasets for training smaller, task-specialized models such as named entity recognition. While this method yields effective models, it assumes that the synthetic dataset is correctly annotated. In this work, we find that (i) current fine-tuning processes simply ignore LLM-introduced annotation noise, resulting in degraded performance and (ii) existing noise-robust losses are not transferable to sequence labeling because annotation noise in named entity recognition is heterogeneous: for example, missing mentions and type errors affect the training signal in different ways. Treating all noisy tokens equally in noise-robust losses and applying a single reweighing criterion for all may therefore remove useful supervision or reinforce incorrect labels. To address this limitation, we propose error-type-aware loss reweighting for NER, which introduces separate reweighing rules for different types of potentially erroneous tokens. Our approach is simple and efficient, does not require additional training resources, and improves F1 by 0.8 - 2.0 percentage points on dataset-level average for noise levels between 15% and 40%, with a maximum improvement of 4.6 percentage points with 24.1% noise on Wikigold.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑