2026年英语词义消歧:当标签成为瓶颈
English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck
- Glite
- Penny Hands Editorial Services(Penny Hands 编辑服务)
- The University of Sheffield(谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究指出英语全词词义消歧的瓶颈已从模型转向标签,通过发布lexEN基准和SenseBench框架,并利用前沿模型重标注语料提升模型性能,同时发现细粒度词义标注存在专家分歧,粗粒度化可同时提升一致性与准确率。
AI中文摘要:
在英语全词词义消歧(WSD)中,标签而非模型已成为瓶颈:前沿大语言模型已足够准确,以至于黄金标准中残留的错误决定了基准排名的走向——这既体现在我们评分的测试集中,也体现在(如我们通过因果分析所示)我们训练的语料中。我们发布了lexEN,一个WSD评估基准,作为对Maru2022的ALL_NEW基准的保守、人工裁决修正层(修改了211个标签,移除了56个),以及SenseBench,一个可审计的LLM WSD评估框架和实时排行榜(57个模型,192次运行)。该任务为受词表约束的多选题(模型从提供的WordNet词义中选择),因此报告的准确率是模型在没有该帮助时所能达到的上限。在lexEN-v1上,前沿LLM的准确率收敛于95%附近(最佳为95.6%),前三大家族在统计上无法区分,且准确率与推理努力和成本在约2500倍的价格跨度内进行权衡。使用前沿模型重新标注SemCor,并在不改变模型的情况下重新训练BEM、ESCHER和ConSeC,在未触及的测试集上提升了数个F1点;我们发布了重新标注的语料库和Glite LENS,一个基于修复后标签训练的298M双编码器(据我们所知,这是报告的最强结果:Raganato ALL 83.6,Maru ALL_NEW 87.4),服务成本约为每百万条0.13美元。在困难条目上,细粒度WordNet词义即使对专家而言也部分不适定(三位评审者的Fleiss kappa=0.537);粗粒度化同时提高了标注者一致性和模型准确率(跨四个词表),使得顶级模型在粗粒度下进入专家一致性区间(在四个词表中的三个上统计等价),但在细粒度下显著低于该区间。当前的约束瓶颈是成本。
英文摘要:
In English all-words word sense disambiguation (WSD), the labels, not the models, have become the bottleneck: frontier LLMs are accurate enough that the errors surviving in the gold standard decide benchmark rankings -- in the test sets we score on and, as we show causally, in the corpus we train on. We release lexEN, a WSD evaluation benchmark built as a conservative, human-adjudicated correction layer over Maru2022's ALL_NEW benchmark (211 labels changed, 56 removed), and SenseBench, an auditable LLM WSD evaluation harness and living leaderboard (57 models, 192 runs). The task is inventory-constrained multiple choice (the model picks from the supplied WordNet senses), so the reported accuracies are a ceiling on what models achieve without that help. On lexEN-v1 the frontier LLMs converge near 95% (best, 95.6%), the top three families are statistically indistinguishable, and accuracy trades off against reasoning effort and cost across a ~2,500x price span. Relabeling SemCor with frontier models and retraining BEM, ESCHER, and ConSeC unchanged lifts them by several F1 points on test sets the relabeling never touched; we release the relabeled corpora and Glite LENS, a 298M bi-encoder trained on the repaired labels -- to our knowledge the strongest reported (83.6 Raganato ALL, 87.4 Maru ALL_NEW) -- serving at ~$0.13 per million items. On hard items, fine-grained WordNet senses are partly ill-posed even for experts (three-reviewer Fleiss kappa=0.537); coarsening raises annotator agreement and model accuracy together across four inventories, placing a top model inside the expert agreement band at coarse granularity (statistically equivalent under three of four) but significantly below it at fine. The binding constraint is now cost.