AI 中文总结
本研究开发乳腺专用病理基础模型CorePath,经7901对数据微调后,在多队列及基准上的CNB诊断任务中性能优于现有模型,结合风险控制模块实现低幻觉的可靠报告生成。
AI 中文摘要
乳腺核心针活检(CNB)是乳腺癌诊断的核心手段,但仍面临挑战:有限的组织采样、病变异质性及细微的形态学重叠会模糊亚型区分。我们开发了CorePath,这是一种乳腺专用多模态病理基础模型,基于PRISM进行微调,使用来自两个中心的7901对CNB全切片图像和诊断报告。在六个CNB队列和两个公开乳腺病理基准上评估时,无需特定任务重新训练,CorePath在癌症检测、侵袭性评估和组织学分型方面始终优于PRISM。它在私人中心的五类CNB组织学分型中,加权受试者工作特征曲线下面积(AUC)达到0.9526至0.9735。在公开基准上,CorePath优于领先的病理基础模型,在BCNB侵袭癌分型中加权AUC为0.7780,在BRACS病变分层中为0.8178,在BRACS细粒度分类中为0.8252,均为最高值。在报告生成方面,CorePath将非乳腺相关幻觉从30.1%降至2.8%,显示出乳腺专用适配后领域保真度的提升。CorePath-CRG进一步将共形亚型置信度门控与“学习后测试”风险控制相结合,实现选择性报告发布、亚型级 fallback 和弃权(不执行)。CorePath-CRG在发布的输出中实现零非乳腺相关幻觉,在病理学家验证的基于LLM的评估分数和跨多数中心的定量报告生成指标中表现出最强的整体性能。这些结果表明,具有统计风险控制的领域专用基础模型为准确的乳腺CNB诊断和可靠的报告生成提供了有前景的方法。
英文摘要
Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breast pathology benchmarks without task-specific retraining, CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping. It achieved weighted area under the receiver operating characteristic curves (AUCs) of 0.9526-0.9735 for five-class CNB histological subtyping across private centers. On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping, 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification. In report generation, CorePath reduced the overall non-breast hallucinations from 30.1% to 2.8%, demonstrating improved domain fidelity after breast-specific adaptation. CorePath-CRG further combined conformal filtering of subtype and binary cancer status predictions with Learn-Then-Test-based threshold calibration to support selective narrative release, diagnostic fallback, and deferral. CorePath-CRG achieved zero non-breast hallucinations among released outputs and showed the strongest overall performance in pathologist-validated LLM-based Evaluation Scores and quantitative report-generation metrics across most centers. These results demonstrate that domain-specialized foundation models with statistical risk control offer a promising approach for accurate breast CNB diagnosis and reliable report generation.
CommentsThe code will be made publicly available upon publication