AI 中文总结
BELXTR提出一种基于多向量晚期交互架构的嵌入模型,通过标记级匹配和主动查询扩展提升生物医学实体链接性能,在多个语料库上平均提升5个百分点recall@1,尤其擅长跨物种基因消歧。
AI 中文摘要
生物医学实体链接将文本中的提及消歧到知识库(KB)中的实体,使其成为信息抽取流程的基石。虽然基于嵌入的模型是此类任务的流行方法,但它们存在一个关键局限性:它们将提及(和实体)压缩为单个向量,迫使模型平均掉关键细粒度差异。我们提出BELXTR,一种基于多向量(又称晚期交互)架构的新型嵌入模型,该模型能够利用标记级匹配信息。BELXTR通过整合现有的任务特定训练目标并探索主动查询扩展,将原始XTR模型扩展到生物医学实体链接。在十个语料库和五个知识库上的实验表明,BELXTR在半数语料库上优于当前最先进水平,平均改进5个百分点(recall@1)。最大的增益出现在具有挑战性的跨物种基因消歧子任务上,其中BELXTR超越了基于LLM的检索-重排流水线,并接近专门的基于规则的系统。我们的结果凸显了多向量模型作为难以维护的基于规则系统或基于LLM的重排在PubMed规模挖掘等场景中成本过高的实用替代方案。复现我们实验的代码可在此https URL找到。
英文摘要
Biomedical Entity Linking disambiguates mentions to entities in a knowledge base (KB), making it the cornerstone of information extraction pipelines. While embedding-based models are a popular approach for the task, they suffer from a key limitation. They compress mentions (and entities) into a single vector, forcing the model to average away crucial fine-grained differences. We present BELXTR, a novel embedding model based on the multi-vector (a.k.a. late interaction) architecture, which allows to leverage token-level matching information. BELXTR extends the original XTR model to biomedical entity linking by integrating an existing task-specific training objective and exploring active query expansion. Experiments across ten corpora and five KBs show that BELXTR improves upon current state-of-the-art in half of the corpora with an average improvement of 5pp recall@1. The largest gains are reported on the challenging cross-species gene disambiguation subtask, where BELXTR outperforms an LLM-powered retrieve-and-rerank pipeline and closely approaches a specialized rule-based system. Our results highlight multi-vector models as a practical alternative to hard-to-maintain rule-based systems or in scenarios where LLM-based reranking is too costly as in PubMed-scale mining. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belxtr.