arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿大语言模型在符合指南且针对具体病例的肿瘤决策中的集体能力边界

A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

Zhang Sheng, Jinming Li, Wangyang Chen, Zhiwei Bao, Yu YoSean Wang

arXiv 2608.28592首次发表:更新:

发表机构

City University of Hong Kong; Fudan University Shanghai Cancer Center; Affiliated Hangzhou First People’s Hospital, Westlake University School of Medicine; Zhejiang University; Shaoxing Vocational & Technical College(香港城市大学; 复旦大学附属肿瘤医院; 西湖大学医学院附属杭州市第一人民医院; 浙江大学; 绍兴职业技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建肿瘤决策边界基准评估9个前沿LLMs,发现其存在临床元判断盲点,需架构干预,模型质量非临床部署主瓶颈,需可检测自身能力边界并路由决策至临床医生的架构。

AI 中文摘要

大语言模型(LLMs)在医学知识考试中取得高分,但现实世界的肿瘤学并非知识测试——它是一系列指南路径选择、升级判断和不确定性下的承诺。现有基准主要衡量事实回忆,而前沿LLMs是否存在无法通过组合模型解决的决策路径盲点仍不明确。我们构建了肿瘤决策边界基准(ODBB),涵盖NCCN指南和结直肠癌病例中的2005个肿瘤决策点,并评估了2025年6月至2026年4月发布的9个前沿LLMs(4个闭源、5个开源权重系列)。一个完全确定性评分器(零LLM推理)将输出分为14种失败类型,由两名肿瘤医生对225项分层样本独立验证(Cohen加权κ分别为0.939和0.790)。将9个模型视为合并超级模型,42.1%(Wilson 95%置信区间40.0--44.3%)的所有项目——1586个NCCN项目中的35.7%和419个结直肠癌病例中的66.4%——未被任何模型答对,失败集中在在任何推理前选择指南路径之间:这是临床元判断中一致的盲点,可能需要架构干预而非更多训练数据。两个针对果断性调优的模型(GPT-5.5、Gemini 3.1 Pro Preview)的不安全承诺频率是7个谨慎模型的3至5倍,且得分未更高。在3--9%的项目中,模型陈述了正确的下一步临床步骤但未作出承诺——这是决策失败,而非知识失败。模型质量不再是临床LLM部署的主要瓶颈;关键约束是认为任何单一模型可作为临床决策唯一基础的假设。进展需要能检测模型达到能力边界并将决策路由给临床医生的架构。

英文摘要

Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen's weighted $κ$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0--44.3%) of all items--35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases--were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3--9% of items, models stated the correct next clinical step yet did not commit to it--failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑