arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27470cs.CLcs.AIcs.DB

选择,而非训练:基于大语言模型(LLM)选择的模块化实体消歧的优势

Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

Fina Polat, Daniel Daza, Pengyu Zhang, Klim Zaporojets, Paul Groth

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出将实体消歧的检索与选择解耦,基于LLM选择器对比不同检索策略,发现无需训练的BM25搭配LLM选择器在ZELDA基准上达到先进性能,还支持检测到检索失败时弃权,提升了实体消歧效果。

中文摘要 AI 辅助

实体消歧(Entity Disambiguation, ED)是构建和使用知识图谱的关键任务。当前最先进的神经方法通常将ED建模为单一任务,尽管它包含两个截然不同的子问题:检索候选实体和在给定上下文的情况下选择正确的实体。双编码器模型在共享嵌入空间中同时优化这两个子任务,迫使表示在高召回率检索与细粒度选择之间进行权衡,并且它们需要训练好的检索器,随着知识图谱的变化,维护成本很高。尽管最近的工作已开始将检索器与基于LLM的选择器相结合,但两个阶段之间的相互作用尚未得到系统研究。在本文中,我们在共享的基于LLM的选择阶段下,对候选生成的检索策略进行了系统比较,将稀疏检索(BM25)、Web KB搜索以及最先进的训练密集检索器与多种开源和闭源LLM相结合。我们表明,一旦将选择任务委托给强大的LLM,训练检索器仅能提供适度的额外价值:完全无需训练的BM25检索器与LLM选择器配对,在ZELDA基准上达到了新的最先进水平,将inKB微F1从82.3提高到86.3(+4);将相同的LLM与训练好的密集检索器配对达到88.5。将检索与选择解耦也揭示了当前ED系统的一个局限性:当检索到的候选中缺少正确实体时,它们被迫预测一个不正确的实体。相比之下,我们的框架在检测到检索失败时允许弃权(不执行)。在奖励正确弃权的评估设置中,无需训练的BM25 + LLM管道达到90.7的F1值。

英文摘要

Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs. State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one given context. Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and they require trained retrievers, which are costly to maintain as knowledge graphs change. While recent work has begun to combine retrievers with LLM-based selectors, the interplay between the two stages has not been studied systematically. In this paper, we present a systematic comparison of retrieval strategies for candidate generation under a shared LLM-based selection stage, combining sparse retrieval (BM25), Web KB search, and a state-of-the-art trained dense retriever with several open- and closed-source LLMs. We show that, once selection is delegated to a capable LLM, training the retriever provides only modest additional value: a fully training-free BM25 retriever paired with an LLM selector reaches a new state of the art on the ZELDA benchmark, raising inKB micro-F1 from 82.3 to 86.3 (+4); pairing the same LLM with a trained dense retriever reaches 88.5. Decoupling retrieval from selection also exposes a limitation of current ED systems: when the correct entity is missing from retrieved candidates, they are forced to predict an incorrect entity. In contrast, our framework allows for abstention when retrieval failure is detected. In an evaluation setting that rewards correct abstentions, the training-free BM25 + LLM pipeline reaches 90.7 F1.

发表机构

  • University of Amsterdam(阿姆斯特丹大学)
  • Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
  • Aarhus University(奥胡斯大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑