发表机构
Institute of Imaging and Computer Vision, RWTH Aachen University; Heinrich Heine University Düsseldorf; Institute of Medical Sciences, Kermanshah University(亚琛工业大学成像与计算机视觉研究所; 杜塞尔多夫海因里希·海涅大学; 克尔曼沙阿大学医学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对透明细胞肾细胞癌分级问题,提出语义引导多模态预处理方法,融合细胞核分类图与RGB图像,使平衡准确率达0.916,优于现有方法,具备良好鲁棒性。
AI 中文摘要
透明细胞肾细胞癌(CCRCC)分级对治疗规划至关重要,但现有方法要么直接分析斑块级图像,要么仅聚焦于细胞核级分类,未与最终肿瘤分级建立关联。我们提出一种语义引导多模态预处理方法,将现有预训练模型生成的细胞核分类图与RGB组织病理学图像相融合,用于基于视觉Transformer(ViT)的CCRCC分级。该方法采用分类图通道拼接与乘法调制,并优化叠加方式以利用细胞核分级信息,同时保留RGB纹理特征。对多种预处理策略的评估显示,语义引导增强的平衡准确率达0.916,优于仅RGB基线方法(0.707)及现有研究的最大投票聚合方法(0.427)。敏感性分析表明,即使在模拟扰动率与当前最先进细胞核分类模型误差阈值匹配的情况下,该方法仍保持比基线高21个百分点的性能提升,说明其具备有效的语义利用能力与实际鲁棒性。这些发现表明,基于预处理的多模态融合可利用现有不完善细胞核分类器的诊断潜力,有效连接此前孤立的细粒度细胞核级分析与粗粒度ViT斑块分类。各分级的类别召回率一致(0.93、0.91、0.91),说明性能提升未集中于多数类。由于敏感性分析扰动的是真实标签图而非实际细胞核模型的预测结果,该结果表征的是模拟误差下的鲁棒性,而非实际上游模型部署时的情况,后者仍为未来工作。
英文摘要
Clear cell renal cell carcinoma (CCRCC) grading is essential for treatment planning, yet existing approaches either analyze patch-level images directly or focus solely on nuclei-level classification, without linking to final tumor grading. We propose a semantic-guided multimodal preprocessing method that integrates nuclei classification maps from existing pre-trained models with RGB histopathology images for Vision Transformer (ViT)-based CCRCC grading. Our approach employs classification map channel concatenation and multiplicative modulation, with optimized overlays to leverage nuclei grading information, while preserving RGB textural features. Evaluation of multiple preprocessing strategies demonstrates that semantic-guided enhancement achieves 0.916 balanced accuracy, outperforming RGB-only baseline (0.707) and max-voting aggregation from prior studies (0.427). Sensitivity analysis reveals that this 21 percentage point improvement over baseline persists even under simulated perturbation at rates matching current state-of-the-art nuclei classification model error thresholds, suggesting both effective semantic utilization and practical robustness. These findings show that preprocessing-based multimodal fusion can leverage the diagnostic potential of existing imperfect nuclei classifiers, effectively bridging previously isolated fine-grained nuclear-level analysis with coarse-grained ViT-based patch classification. Per-class recall was consistent across grades (0.93, 0.91, 0.91), indicating that gains are not concentrated in the majority class. Because the sensitivity analysis perturbs ground-truth maps rather than predictions from an actual nuclei model, this result characterizes robustness under simulated error rather than deployment with a real upstream model, which remains for future work.
Comments12 pages, 3 Figures, COMPAYL++ MICCAI 2026