发表机构
PICT; L3Cube Labs; Indian Institute of Technology Madras(浦那信息技术学院; L3Cube实验室; 印度理工学院马德拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长文档NER中截断和分块导致的实体碎片问题,提出基于重叠滑动窗口的推理流水线,无需重训即可提升文档级F1分数至0.8902且波动低于1%。
AI 中文摘要
命名实体识别(NER)是自然语言处理中的一项核心任务,但基于Transformer的句子级模型由于固定的输入长度限制,在处理长文档时面临困难:截断会丢失内容,而不重叠的分块会在段边界处切碎实体。我们提出了一种仅推理的流水线,通过重叠滑动窗口将冻结的NER模型MahaNER-BERT(在MahaNER语料库上微调)扩展到文档级预测,并将窗口结果合并为单一标注,无需任何重新训练或架构更改。我们在由MahaNER测试集构建的六个文档级语料库上评估该流水线,采用两种策略:Normal Repeat(重复句子序列以延长长度同时保持上下文连续性)和Random Repeat(拼接不同序列以产生更长、异质的输入),每种策略在三个长度级别和多种滑动窗口配置下实例化。模型保持了最高达0.8902的宏F1分数,且无论文档长度或构建策略如何,变化幅度均低于一个百分点。与传统的非窗口方法相比,滑动窗口流水线避免了非重叠分割引入的边界碎片化错误,从而产生了一致更高且更稳定的文档级F1分数。
英文摘要
Named Entity Recognition (NER) is a core NLP task, but transformer-based sentence-level models struggle with long documents because of fixed input-length limits: truncation drops content, and non-overlapping chunking fragments entities at segment boundaries. We introduce an inference-only pipeline that extends a frozen NER model, MahaNER-BERT, fine-tuned on the MahaNER corpus, to document-level prediction via overlapping sliding windows that are merged into a single annotation, without any retraining or architectural change. We evaluate the pipeline on six document-level corpora built from the MahaNER test set using two strategies: Normal Repeat, which duplicates sentence sequences to extend length while preserving contextual continuity, and Random Repeat, which concatenates distinct sequences to produce longer, heterogeneous inputs, each instantiated at three length levels, across several sliding-window configurations. The model retains a macro F1-score of up to 0.8902, with variation staying below one percentage point regardless of document length or construction strategy. Compared with the conventional non-windowed approach, the sliding-window pipeline avoids the boundary-fragmentation errors introduced by non-overlapping segmentation, yielding consistently higher and more stable document-level F1-scores.