arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20593cs.CL

WiC 并非 WSD:关于大语言模型与词汇歧义消解的研究

WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

Yi Zhou, Kiamehr Rezaee, Danushka Bollegala, Mohammad Taher Pilehvar, Jose Camacho-Collados

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过对比WiC与传统WSD任务,发现提供候选词义能提升大语言模型在WiC上的表现,并指出模型错误多源于标签歧义和过度细粒度的词义区分。

中文摘要 AI 辅助

词义消歧(Word-in-Context, WiC)任务对语言模型而言仍具挑战性,尽管在词汇语义任务上近期已有进展。我们假设,这一困难不仅源于比较一个词的两种上下文用法,还源于缺乏一个明确指定语义粒度级别的词义清单。我们在相似设置下评估了开源大语言模型在WiC和传统词义消歧(Word Sense Disambiguation, WSD)任务上的表现。我们发现,提供候选词义(类似于传统WSD的做法)在所有设置下均能提升WiC性能。总体而言,显式的词义信息有助于模型做出更一致、更具针对性的判断。人工评估进一步表明,许多表面上的WiC错误反映了标签歧义或模型与标注者之间词义边界的不匹配,而非简单的词汇理解失败。特别是,结果显示大语言模型会过度思考词义区分,常常因过于细粒度的区分而导致错误。

英文摘要

Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory that specifies the relevant level of semantic granularity. We evaluate open LLMs on WiC and traditional Word Sense Disambiguation (WSD) under similar settings. We find that providing candidate senses, similar to what is done in traditional WSD, improves WiC performance in all settings. In general, explicit sense information helps models make more consistent and targeted judgements. Human evaluation further shows that many apparent WiC errors reflect label ambiguity or mismatches between model and annotator sense boundaries rather than simple failures of lexical understanding. In particular, results show that LLMs overthink the sense distinction often leading to errors based on overly fine-grained distinctions.

发表机构

  • University of Liverpool(利物浦大学)
  • Amazon(亚马逊)
  • Cardiff University(卡迪夫大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑