发表机构
Xunfei Healthcare Technology Co., Ltd.(讯飞医疗科技股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SEER-Bench基准,对比四种监督形式后发现EMQ在相同预算下可提升大型语言模型的医学知识更新效果,4B模型在该监督下于SEER-Bench取得64.8%答案准确率,证明医学知识更新与监督形式相关。
AI 中文摘要
医学知识不断变化,这使得大型语言模型容易依赖过时但临床合理的信息。我们在匹配训练预算的设置下研究监督形式是否影响医学知识更新。我们推出SEER-Bench,这是一个基于最新版本SEER研究数据发布的时间锚定肿瘤分期基准,并将NCCN肿瘤指南中的相同医学更新事件转换为四种监督形式:EMQ、MSQ、FITB和SAQ。在SEER-Bench和HealthBench Professional上,在相同预算的SFT变体中,EMQ实现了最稳定的外部迁移和保留。在EMQ监督下,更新后的4B模型在时间锚定的肿瘤分期上取得了有竞争力的结果,在SEER-Bench上达到64.8%的答案准确率和59.6%的原理准确率。诊断分析表明,EMQ暴露了更密集的临床对比信号,同时在保持与基础模型更小偏差的情况下保留了判别性表示。这些结果表明,医学知识更新不仅取决于更新算法,还取决于知识被构建为监督的方式。
英文摘要
Medical knowledge changes continually, making large language models vulnerable to relying on outdated yet clinically plausible information. We study whether the format of supervision affects medical knowledge updating under a matched training-budget setting. We introduce SEER-Bench, a temporally anchored oncology-staging benchmark curated from the latest versioned SEER Research Data release, and render identical medical update events from NCCN oncology guidelines into four supervision formats: EMQ, MSQ, FITB, and SAQ. Across SEER-Bench and HealthBench Professional, EMQ gives the most stable external transfer and retention among same-budget SFT variants. With EMQ supervision, the updated 4B model produces competitive results on temporally anchored oncology staging, reaching 64.8% answer accuracy and 59.6% rationale accuracy on SEER-Bench. Diagnostic analyses suggest that EMQ exposes denser clinical contrast signals while preserving discriminative representations with smaller movement from the base model. These results show that medical knowledge updating depends not only on the update algorithm, but also on how knowledge is structured as supervision.
CommentsAccepted to Findings of EMNLP 2026