arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03154cs.CL

ANCHOR-RE:用于接地生物医学关系抽取的智能体神经符号框架

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Shufan Ming, Yikun Han, Gibong Hong, Rui Zhang, Halil Kilicoglu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ANCHOR-RE框架,将本体引导推理等集成到LLM推理中,在多个BioRE基准及2026年生物医学文章数据集上,该框架性能优于直接LLM提示及部分现有方法,是实用的无训练生物医学文献挖掘方法。

中文摘要 AI 辅助

生物医学关系抽取(BioRE)从生物医学文献中提取结构化知识,用于知识库构建、假设生成等应用。传统符号系统如SemRep具有高准确率但召回率有限,而大型语言模型(LLM)的上下文推理能力更强,但仍易出现假阳性预测。我们开发了ANCHOR-RE,该框架将本体引导推理、外部知识接地和数据驱动验证规则集成到LLM推理中。我们在三个BioRE基准数据集(SemRepGS、DDI和ChemProt)上,使用专有和开放权重LLM对其进行评估。为评估基准数据集之外的泛化能力,同时减少LLM预训练污染带来的潜在评估偏差,我们使用100篇2026年发表的生物医学文章进行时间评估。使用专有主干模型时,ANCHOR-RE的表现优于直接LLM提示,在SemRepGS上的微F1从0.654提升至0.676,在DDI上从0.769提升至0.872,在ChemProt上从0.939提升至0.941。在DDI和ChemProt上,它还优于之前报告的仅推理方法,且无需参数更新即可接近微调或指令微调系统。开放权重LLM也观察到类似的性能提升,表明该优势不限于专有主干模型。在截止后数据集上,对500个随机抽样预测的人工评估得出准确率为69%,在之前未见过的生物医学文献上保持了一致的准确率。神经符号推理可在不进行微调的情况下提高基于LLM的BioRE的可靠性。跨多个基准、模型家族和截止后文献的结果支持ANCHOR-RE作为一种实用的无训练生物医学文献挖掘方法。

英文摘要

Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep provide high precision but limited recall, while large language models (LLMs) offer stronger contextual reasoning but remain prone to false-positive predictions. We developed ANCHOR-RE, a framework that integrates ontology-guided reasoning, external knowledge grounding, and data-driven verification rules into LLM inference. We evaluated it on three BioRE benchmarks (SemRepGS, DDI, and ChemProt) using both proprietary and open-weight LLMs. To assess generalizability beyond benchmark datasets while reducing potential evaluation bias from LLM pretraining contamination, we conducted a temporal evaluation using 100 biomedical articles published in 2026. With the proprietary backbone, ANCHOR-RE outperformed direct LLM prompting, improving micro-F1 from 0.654 to 0.676 on SemRepGS, from 0.769 to 0.872 on DDI, and from 0.939 to 0.941 on ChemProt. On DDI and ChemProt, it also outperformed previously reported inference-only methods and approached fine-tuned or instruction-tuned systems without parameter updates. Similar performance gains observed with open-weight LLMs indicate that the benefits were not limited to the proprietary backbone. On the post-cutoff set, manual assessment of 500 randomly sampled predictions yielded a precision of 69%, maintaining consistent precision on previously unseen biomedical literature. Neuro-symbolic reasoning can improve the reliability of LLM-based BioRE without fine-tuning. Results across multiple benchmarks, model families, and post-cutoff literature support ANCHOR-RE as a practical training-free approach to biomedical literature mining.

补充信息

↑