MUFASA:面向计算病理学中可靠模型推理的信息效用感知预处理框架
MUFASA: An Information Utility-Aware Preprocessing Framework for Reliable Model Reasoning in Computational Pathology
浏览论文内容
中文总结 AI 辅助
本研究提出MUFASA框架,通过整合伪影掩蔽、图像块过滤等技术优化计算病理学WSI预处理,提升下游AI模型在多癌种任务中的性能与可靠性,揭示预处理对模型有效性的关键作用。
中文摘要 AI 辅助
可靠的计算病理学依赖于预处理方法,这些方法能识别全切片图像(WSI)中的信息丰富的组织区域,同时排除伪影和低效用区域。然而,现有的预处理流程常保留这类区域,或丢弃与诊断相关的组织,从而限制了下游模型在异质性队列中的性能、可靠性和鲁棒性。本文中,我们系统评估了这些区域对多个临床相关应用中下游AI模型性能的影响,并提出了MUFASA——一种适用于HE染色WSI的可泛化信息效用感知预处理框架,该框架可排除伪影和低效用区域,同时保留具有生物学意义的组织。MUFASA整合了切片级伪影掩蔽、染色感知的图像块过滤、基于重建的图像块效用分层,以及对早期阶段过度过滤的组织图像块的针对性恢复。在不同癌症队列的肿瘤诊断、肿瘤亚型分类、生物标志物状态预测和生存预后任务中,与广泛使用的预处理基线相比,MUFASA始终提升了下游模型的性能。这些提升伴随模型热图中伪影相关归因的减少,表明保留的组织与模型注意力之间的对齐得到改善。我们的研究结果确立了WSI预处理是下游模型性能和有效性的关键决定因素,揭示了即使是准确的预测也可能隐藏重要的失败模式,这些失败模式源于保留的含伪影及低信息效用图像块所导致的解剖学上不可信的推理。
英文摘要
Reliable computational pathology depends on preprocessing methods that identify informative tissue regions while excluding artifacts and low-utility regions from whole-slide images (WSI). However, existing preprocessing pipelines often retain such regions or discard diagnostically relevant tissue, thereby limiting downstream model performance, reliability, and robustness across heterogeneous cohorts. Here, we systematically evaluate how these regions affect downstream AI model performance across multiple clinically relevant applications and introduce MUFASA, a generalizable, information utility-aware preprocessing framework for H&E-stained WSI that excludes artifacts and low-utility regions while preserving biologically meaningful tissue. MUFASA integrates slide-level artifact masking, stain-aware tile filtering, reconstruction-based utility stratification of tiles, and targeted recovery of tissue tiles that are over-filtered by earlier phases. Across tumor diagnosis, tumor subtyping, biomarker status prediction, and survival prognostication tasks in diverse cancer cohorts, MUFASA consistently improves downstream model performance relative to widely used preprocessing baselines. These gains are accompanied by reduced artifact-associated attribution in model heatmaps, indicating improved alignment between retained tissue and model attention. Our findings establish WSI preprocessing as a critical determinant of downstream model performance and validity, revealing that even accurate predictions can conceal important failure modes stemming from anatomically implausible reasoning driven by retained artifact-containing and low information-utility tiles.
发表机构
- Stanford University School of Medicine(斯坦福大学医学院)
- CHA Gangnam Medical Center, CHA University School of Medicine(CHA江南医疗中心,CHA大学医学院)
机构由 AI 辅助整理,请以论文原文为准。