发表机构
Queen’s University Belfast; University of Texas at Austin(贝尔法斯特女王大学; 德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对生物医学声明验证,在CARE-XAI基准上对比多种模型,发现微调大型语言模型生成证据能力最强,PubMed检索效果因来源而异,还提出Bio-GRACE工具评估检索效用,推动选择性检索发展。
AI 中文摘要
生物医学事实核查系统不仅要预测一项声明是得到支持、被反驳还是未被提及,还应生成忠实、完整且对核查有用的证据。我们在涵盖五个生物医学和健康事实核查来源的统一基准CARE-XAI上研究这种证据生成设置,在共享评估协议下比较基础指令大型语言模型、PubMed检索增强型大型语言模型、微调后的大型语言模型、仅标签型大型语言模型和生物医学编码器分类器。生物医学分类器在仅判断结果预测上仍表现最强,而微调后的大型语言模型是生成证据能力最强的系统。PubMed检索的效果好坏参半:它对PubMedQA和SciFact等与PubMed对齐的来源有帮助,但会在更广泛的公共卫生声明上干扰模型。我们引入Bio-GRACE,这是一种金标准参考归一化诊断工具,用于衡量检索到的证据是否恢复了参考证据的决策益处。Bio-GRACE表明检索效用具有来源依赖性,推动了选择性检索的发展,并揭示了为何检索召回率和词汇证据重叠不足以用于生物医学事实核查。
英文摘要
Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence that is faithful, complete, and useful for verification. We study this evidence-generation setting on CARE-XAI, a unified benchmark spanning five biomedical and health fact-checking sources. We compare base instruction LLMs, PubMed retrieval-augmented LLMs, fine-tuned LLMs, label-only LLMs, and biomedical encoder classifiers under a shared evaluation protocol. Biomedical classifiers remain strongest for verdict-only prediction, while fine-tuned LLMs are the strongest evidence-generating systems. PubMed retrieval is mixed: it helps PubMed-aligned sources such as PubMedQA and SciFact, but can distract models on broader public-health claims. We introduce Bio-GRACE, a gold-reference-normalized diagnostic for measuring whether retrieved evidence recovers the decision benefit of reference evidence. Bio-GRACE shows that retrieval utility is source-dependent, motivates selective retrieval, and exposes why retrieval recall and lexical evidence overlap are insufficient for biomedical fact-checking.