OpenTumorBoard:多学科肿瘤委员会讨论轨迹的真实世界基准
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
AI总结:
OpenTumorBoard是一个基于真实肿瘤委员会讨论轨迹的基准,含611个病例和19,157轮讨论,评估14个LLM,发现性能有限,但微调和强化学习可提升,支持多学科癌症决策模型开发。
AI中文摘要:
多学科肿瘤委员会通过专家讨论整合多模态临床观察和纵向患者病史,然而基准测试很少捕捉这些真实世界的轨迹。我们引入了OpenTumorBoard,一个包含611个患者病例和19,157轮讨论的基准,涉及十个专家角色,这些讨论转录自YouTube上公开可用的12,534分钟肿瘤委员会录音。该基准评估两种设置:专家轮次,其中大型语言模型(LLM)回应真实讨论中提出的临床重要问题;以及委员会模拟,其中模型生成完整的来回讨论,并就治疗建议、手术计划、后续行动和临床试验匹配达成共识。对14个通用前沿和医学LLM的评估揭示了显著局限性:最佳模型在临床等效性上得分为3.43(满分5分),在与记录的委员会结论一致性上得分为2.78(满分5分)。监督微调和强化学习在保留测试集上提高了性能,表明真实世界的讨论轨迹可以支持模型适应。三位医学博士专家审查了基准的一个子集,发现患者病例的信息覆盖率和事实性高,提取的共识结论保真度强。我们将发布OpenTumorBoard及其自动化整理流程,以支持用于多学科、个性化癌症决策的LLM开发和评估。
英文摘要:
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly available tumor board recordings on YouTube. The benchmark evaluates two settings: SPECIALIST TURN, in which an LLM responds to a clinically significant question posed during a real discussion, and BOARD SIMULATION, in which it generates an entire back-and-forth discussion and reaches a consensus on therapy recommendations, surgical plans, next actions and clinical trial matching. Evaluation of 14 general-purpose frontier and medical LLMs reveals substantial limitations: the best models score 3.43 out of 5 in clinical equivalence to specialist answers and 2.78 out of 5 in alignment with recorded board conclusions. Supervised finetuning and reinforcement learning improve performance on a held-out test set, suggesting that real-world discussion trajectories can support model adaptation. Three M.D. experts review a subset of the benchmark, finding high information coverage and factuality of patient cases and strong fidelity of extracted consensus conclusions. We will release OpenTumorBoard and its automated curation pipeline to support the development and evaluation of LLMs for multidisciplinary, personalized cancer decision-making.