arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MedKIT:评估大型语言模型中的知识整合与泛化能力

MedKIT: Evaluating Knowledge Integration and Generalization in Large Language Models

Lukas Thede, Yash Kumar Atri, David Chen, Danielle Bitterman, Matthias Bethge, Tom Hartvigsen, Zeynep Akata

arXiv 2609.38543首次发表:更新:

发表机构

University of Tübingen; Tübingen AI Center; Helmholtz Munich; Munich Center for Machine Learning (MCML); University of Virginia; University of British Columbia; Harvard Medical School; Technical University of Munich(蒂宾根大学; 蒂宾根人工智能中心; 亥姆霍兹慕尼黑中心; 慕尼黑机器学习中心; 弗吉尼亚大学; 不列颠哥伦比亚大学; 哈佛医学院; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对医学知识更新中模型整合知识后实际可用性不足的问题,提出MedKIT基准,通过细粒度探针评估12种策略在5种模型上的表现,发现回忆与可用知识存在差距,为知识整合方法开发提供测试平台。

AI 中文摘要

不断演变的现实世界知识要求模型持续更新。尤其在医学领域,随着临床证据随时间变化,过时的知识可能带来安全风险。现有的知识整合评估侧重于事实回忆,对于新整合的知识是否真正可用提供的见解有限。我们的基准MedKIT(医学知识整合与迁移)在现实的临床更新序列下,对模型如何整合和应用知识进行了细粒度评估。每个实例对应一个源自临床证据的事实更新,并配有针对性的探针,以评估跨词汇变化、关系转换、组合推理和开放式操作化的迁移能力,以及用于知识保留的局部性测试。利用MedKIT,我们对5种不同模型(包括通用型和医学专用LLM)的12种知识整合策略进行了大规模实证研究。我们的结果揭示了回忆与可用知识之间的一致差距:虽然大多数方法在原始更新任务和词汇变化下取得了显著提升,但关系泛化能力有限,且没有任何方法在组合或操作任务上产生有意义的改进。这些发现凸显了知识整合中的一个根本性挑战,并将MedKIT定位为一个测试平台,用于开发使新整合的知识在跨任务和情境中更一致可用的方法。

英文摘要

Constantly evolving real-world knowledge necessitates models to be updated continuously. Especially in medicine, as clinical evidence changes over time, outdated knowledge can pose safety risks. Existing evaluations of knowledge integration focus on factual recall, offering limited insight into whether newly integrated knowledge is actually usable. Our benchmark MedKIT (Medical Knowledge Integration and Transfer) provides a granular evaluation of how models integrate and apply knowledge under realistic sequences of clinical updates. Each instance corresponds to a factual update derived from clinical evidence, paired with targeted probes that assess transfer across lexical variation, relational transformations, compositional reasoning, and open-ended operationalization, as well as locality tests for knowledge preservation. Using MedKIT, we conduct a large-scale empirical study of 12 knowledge integration strategies across 5 diverse models, including both general-purpose and medical LLMs. Our results reveal a consistent gap between recall and usable knowledge: while most methods achieve strong gains on the original update task and under lexical variation, relational generalization is limited, and no method yields meaningful improvements on compositional or operational tasks. These findings highlight a fundamental challenge in knowledge integration and position MedKIT as a testbed for developing methods that make newly integrated knowledge more consistently usable across tasks and contexts.

CommentsAccepted at NeurIPS 2026 (Evaluations & Datasets Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑