端侧命名实体识别:准确性、成本、可靠性与置信度的可部署性研究
On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence
浏览论文内容
中文总结 AI 辅助
本研究在端侧部署场景下系统比较了九种NER模型,发现编码器模型在规模、延迟和输出有效性上占优,而生成式LLM在准确性上具有竞争力,并刻画了置信度校准特性。
中文摘要 AI 辅助
命名实体识别(NER)越来越需要在设备端(无API、低延迟、数据保存在本地)运行。实践者的问题不在于排行榜,而在于哪种模型可以部署、如何在没有人工标注的情况下评估它,以及其置信度是否可信。我们共同回答了这些问题。我们将九个系统置于三种范式和13M至8B参数范围内:一个经典标注器(spaCy)、双向编码器专家(GLiNER,166至460M)以及本地运行的生成式LLM(Qwen3-0.6B/1.7B/4B-Instruct,DeepSeek-R1-1.5B/8B),在三个特征各异的数据集上进行测试,并报告准确性以及文献中遗漏的两个维度:延迟和输出有效性。由于我们的语料库(RSS-News)没有金标准,我们从跨家族LLM评审团构建了银标准,然后测量其相对于基准金标准和语料库完整人工再验证的保真度(严格F1为0.95,由于人工金标准是银标准种子化的,这是一个上限);金标准的来源翻转了范式排名,从LLM生成的银标准转向人工金标准提高了每个编码器的性能并降低了每个生成模型的性能。仅就准确性而言,一个4B的指令微调LLM具有竞争力(在干净的新swire上领先),因此编码器的优势在于可部署性:它以九分之一到二十四分之一的规模、毫秒到秒级的延迟、零格式错误输出,与生成模型匹配或略微落后,而最小的生成模型在长输入上会产生高达27%的无效输出,这种失败通过规模而非输出预算来修复。然后我们刻画了GLiNER的每跨度置信度:它在排序正确性方面表现良好(AUROC 0.76至0.86),但过于自信(ECE 0.24至0.47,通过温度缩放减半);阈值化带来了小幅诚实的样本外F1增益;一个全本地的小到大型级联在成本匹配的随机路由上提供了适度的、依赖语料库的增益;置信度跟踪正确性但不跟踪新颖性。每个数字都从每跨度记录离线重新计算。
英文摘要
Named-entity recognition (NER) is increasingly wanted on-device (no API, low latency, data kept local). The practitioner's question is not the leaderboard but which model is deployable, how to evaluate it without human annotation, and whether its confidence can be trusted. We answer these jointly. We place nine systems across three paradigms and 13 M to 8 B parameters: a classical tagger (spaCy), bidirectional-encoder specialists (GLiNER, 166 to 460 M), and generative LLMs run locally (Qwen3-0.6B/1.7B/4B-Instruct, DeepSeek-R1-1.5B/8B), on three datasets of differing character, and report accuracy plus two axes the literature omits: latency and output validity. Because our corpus (RSS-News) had no gold, we built silver gold from a cross-family LLM judge panel, then measured its fidelity against benchmark gold and a full human re-validation of the corpus (strict F1 0.95, an upper bound since the human gold was silver-seeded); gold provenance flips the paradigm ranking, moving from LLM-authored silver to human gold raises every encoder and lowers every generative model. On accuracy alone a 4 B instruct LLM is competitive (it leads on clean newswire), so the encoder's case is deployability: it matches or slightly trails at one-ninth to one-twenty-fourth the size, at millisecond-to-second latency, with zero malformed output, while the smallest generative models emit up to 27% invalid output on long inputs, a failure fixed by scale, not output budget. We then characterize GLiNER's per-span confidence: it ranks correctness well (AUROC 0.76 to 0.86) but is overconfident (ECE 0.24 to 0.47, halved by temperature scaling); thresholding gives a small honest out-of-sample F1 gain; an all-local small-to-large cascade gives a modest, corpus-dependent gain over cost-matched random routing; and confidence tracks correctness but not novelty. Every number recomputes offline from per-span records.