arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11770cs.CL

医学大语言模型研究中不断扩大的评估差距:2023至2026年

The widening evaluation gap in medical large language model research 2023 to 2026

  • XU Exponential University of Applied Sciences(XU指数应用科学大学)
  • Abu Dhabi University(阿布扎比大学)
  • Zayed University(扎耶德大学)

机构由 AI 辅助整理,请以论文原文为准。

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif

AI总结:

本研究揭示2023至2026年医学大语言模型评估存在扩大滞后,随机试验更倾向评估过时模型,指出严谨性与时效性的矛盾源于模型选择而非研究时间线。

AI中文摘要:

大语言模型每隔几个季度就会被更新换代;而临床证据的生成需要数年时间。我们探讨了医学研究是否跟上了其所评估系统的步伐。在2023年1月至2026年6月期间,PubMed在十四个临床领域返回了11,628条记录,增长了45倍;其中2.5%采用了随机、对照或前瞻性设计。评估滞后——从研究中最新的具名模型发布到其自身发表——从1.33个季度扩大到6.08个季度。由于已停用模型会机械性地老化,我们将其与一个保持模型组成不变的反事实基准进行了比较:向更新系统的迁移仅抵消了56%的漂移(95%置信区间50-65)。随机试验评估的模型中位数比其他设计老4.6个季度(P = 3 x 10^-19),然而在那些点名了仍在开发中的模型的研究中,不同设计之间没有差异;62%的随机试验评估了已停用的模型家族。严谨性与时效性之间存在张力,而这种张力反映的是模型选择问题,而非研究时间线问题。

英文摘要:

Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.

↑