发表机构
Radboud University Medical Center; Erasmus University Medical Center; Oncode Institute; Canisius Wilhelmina Hospital; Tampere University; Tampere University Hospital(拉德堡德大学医学中心; 伊拉斯姆斯大学医学中心; Oncode研究所; 卡尼修斯·威廉米娜医院; 坦佩雷大学; 坦佩雷大学医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对高危非肌层浸润性膀胱癌,CHIMERA挑战赛利用多模态数据(组织病理学、结构化数据和RNA测序)进行应答亚型分类和进展预测,最佳模型分别取得F1 0.73和C指数0.68,并揭示了模态贡献和预测难点。
AI 中文摘要
高危非肌层浸润性膀胱癌(HR-NMIBC)具有显著的复发和进展风险,而当前的临床风险分层仍然有限。CHIMERA作为一个多模态人工智能挑战赛被建立,旨在标准化评估下对HR-NMIBC进行预测基准测试。任务BRS(应答亚型分类)利用组织病理学和结构化临床病理数据预测RNA-seq定义的卡介苗(BCG)应答亚型,而任务Progression(进展预测)则利用组织病理学、结构化数据和RNA测序对进展时间进行建模。一个包含368名患者的多模态数据集被划分为公开训练集和隐藏验证集及测试集。总共收到159份提交,并选出13个表现最佳的模型进行基准测试。最佳模型在任务BRS中实现了0.73的加权F1分数,在任务Progression中实现了0.68的C指数。挑战赛后的分析揭示了任务依赖的模态贡献、队列依赖的性能下降以及对缺失结构化数据的敏感性。在任务BRS中,组织病理学部分补偿了来源于病理学的结构化变量,而进展模型则表现出对互补输入的更大依赖性。跨模型误差分析进一步识别出在不同架构中始终难以预测的患者,其中T1亚分期与预测难度相关。这些发现凸显了可迁移性的障碍以及缺失感知建模和独立多机构验证的重要性。CHIMERA为膀胱癌提供了一个标准化的多模态基准,并为研究不仅限于模型性能,还包括鲁棒性、信息充分性和患者层面预测失败提供了一个框架。
英文摘要
High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification remains limited. CHIMERA was established as a multimodal AI challenge to benchmark prediction in HR-NMIBC under standardized evaluation. Task BRS predicts RNA-seq-defined BCG Response Subtypes from histopathology and structured clinicopathological data, whereas Task Progression models time-to-progression using histopathology, structured data, and RNA sequencing. A multimodal dataset of 368 patients was divided into public training and hidden validation and test sets. In total, 159 submissions were made, and 13 top-performing models were selected for benchmarking. The best models achieved a weighted F1 score of 0.73 for Task BRS and a C-index of 0.68 for Task Progression. Post-challenge analyses revealed task-dependent modality contributions, cohort-dependent performance degradation, and sensitivity to missing structured data. In Task BRS, histopathology partly compensated for pathology-derived structured variables, whereas progression models showed greater dependence on complementary inputs. Cross-model error analysis further identified patients that were consistently difficult across different architectures, with T1 substage associated with prediction difficulty. These findings highlight barriers to transportability and the importance of missingness-aware modeling and independent multi-institutional validation. CHIMERA provides a standardized multimodal benchmark for bladder cancer and a framework for studying not only model performance, but also robustness, information sufficiency, and patient-level prediction failure.
CommentsarXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission.