弥合词汇差异:大语言模型辅助的低成本零样本科学实体链接
Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking
浏览论文内容
中文总结 AI 辅助
针对科学实体链接中词汇重叠不足、零样本方法计算成本高且噪声多的问题,提出Sci-ZSEL框架,结合选择性LLM别名生成、本体感知过滤和伪标注微调,在含低词汇重叠的新动物科学基准及其他基准上表现优于基线。
中文摘要 AI 辅助
科学领域的实体链接(EL)与通用领域的实体链接不同,因为提及项和实体名称往往缺乏词汇重叠。另一个挑战是科学领域使用的专业术语在通用领域预训练的模型中很少遇到,因此在通用领域训练的模型难以迁移到科学领域。为解决这一问题,领域内微调是自然的解决方案,但许多科学领域缺乏专家标注的数据,催生了对零人工标注方法的需求。现有的零样本方法严重依赖大语言模型(LLM)为整个提及语料库生成别名,这会产生大量计算成本,且这些方法没有机制过滤大语言模型产生的噪声。为应对这些挑战,我们提出Sci-ZSEL框架,该框架选择性使用大语言模型生成实体别名以控制计算成本,并应用感知本体的过滤器去除语义漂移到本体邻居的别名。然后,过滤后的别名用于构建伪标注的提及-实体对进行微调。为评估低词汇重叠下的实体链接,我们还发布了一个新的动物科学实体链接基准,该基准链接到三个牲畜性状本体,其中提及项和实体的词汇重叠比现有基准低得多。在五个基准上,Sci-ZSEL优于未微调的基线,在非重叠提及上最有用,将其与精心整理的同义词结合在大多数设置中可获得最佳性能。
英文摘要
Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is that specialized terminology is used in the scientific domain, which is rarely encountered in models pretrained on general domains. Therefore, models trained on general domains transfer poorly to scientific domains. To address this, in-domain fine-tuning is the natural remedy. However, many scientific domains lack expert-annotated data, motivating the need for a zero-human-annotation approach. Existing zero-shot methods heavily rely on LLMs to generate aliases across entire mention corpora, which incurs substantial computational cost, and those methods provide no mechanism to filter out noise from LLMs. To address these challenges, we propose Sci-ZSEL, a framework that selectively generates entity aliases with an LLM to control computational cost, and applies an ontology-aware filter to remove aliases that semantically drift toward ontology neighbors. Then, filtered aliases are used to construct pseudo-labeled mention-entity pairs for fine-tuning. To enable evaluation of EL under low lexical overlap, we also release a new animal science EL benchmark linked to three livestock trait ontologies, where mentions and entities exhibit substantially lower lexical overlap than in existing benchmarks. Across five benchmarks, Sci-ZSEL outperforms the non-fine-tuned baseline, is most useful on nonoverlapping mentions, and combining it with curated synonyms gives the best performance in most settings.
发表机构
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。