发表机构
Stanford University; Dana-Farber Cancer Institute; Harvard Medical School(斯坦福大学; 丹娜-法伯癌症研究所; 哈佛医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM在分子肿瘤委员会中的安全性评估,提出OpenMTB-Audit基准和MTB-AuditAgent框架,解决过度拒绝问题,将过度拒绝率降至6.7%,准确率达91.2%,并揭示临床专家在信息充分性边界上的分歧。
AI 中文摘要
分子肿瘤委员会整合基因组发现、临床背景和治疗证据以支持精准肿瘤学。随着人工智能进入这一工作流程,一个关键的安全挑战是区分真正不支持的推荐与仍需肿瘤学家审查的证据支持选项,后者可能因信息不完整、ECOG体能状态不佳或其他临床注意事项而需要审查。我们推出了OpenMTB-Audit,一个包含500个合成非小细胞肺癌病例的开源基准,涵盖五类对抗性错误类别和四个安全标签:支持、部分支持、不支持和信息不足。在八种大型语言模型配置中,我们发现了普遍的过度拒绝:所有LLM配置在83.3%-100%的真实部分支持病例中未能保留“部分支持”标签,通过标签坍缩而非临床校准推理获得了较高的总体安全分数。为解决这一局限性,我们开发了MTB-AuditAgent,一个确定性的七模块框架,将证据验证、缺失信息检测、安全分类和弃权(不执行)分离。它将过度拒绝率降至6.7%,并达到91.2%的准确率(95%置信区间:88.6%-93.6%)。一项由两位肿瘤学家参与的标注研究发现,分歧集中在信息充分性与治疗优化之间的边界,强调了保留临床上有意义的区分的重要性。
英文摘要
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
CommentsAccepted for oral presentation and publication at the Pacific Symposium on Biocomputing (PSB) 2027