发表机构
Stony Brook University; University of Utah(石溪大学; 犹他大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对病理学报告生成难以衡量进展的问题,提出标准化基准测试和评估框架,用三种编码器在三个数据集上评估四种方法,引入临床报告质量评分(CRQS),揭示模型和编码器差异,为病理学报告生成评估奠定可重复基础。
AI 中文摘要
从全切片图像(WSIs)生成病理学报告是一个快速发展的多模态学习问题,但由于现有研究使用异构数据集、模型设置、视觉编码器和评估协议,进展难以衡量。常用的自然语言生成指标主要奖励词汇相似性,往往无法检测到临床相关错误。我们提出了一个用于病理学报告生成的标准化基准测试和评估框架。该基准测试使用三种病理学基础编码器,在三个数据集上评估四种代表性方法,并标准化了预处理、特征提取、训练、解码和评估。核心贡献是临床报告质量评分(CRQS),实验表明传统语言生成指标与临床正确性的一致性较弱,而CRQS揭示了词汇指标未能捕捉到的模型和编码器之间具有临床意义的差异。该基准测试、公共即插即用框架和CRQS为严格评估病理学报告生成奠定了可重复的基础。
英文摘要
Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure because existing studies use heterogeneous datasets, model settings, visual encoders, and evaluation protocols. Moreover, commonly used natural language generation metrics, including BLEU, ROUGE, and METEOR, primarily reward lexical similarity and often fail to detect clinically consequential errors such as omitted diagnoses, hallucinated findings, or discordant tumor attributes. We present a standardized benchmark and evaluation framework for pathology report generation. The benchmark evaluates four representative methods across three datasets (TCGA, HistAI, and REG 2025) using three pathology foundation encoders (CONCHv1.5, UNI2-h, and H-Optimus-1). Our framework standardizes preprocessing, feature extraction, training, decoding, and evaluation, enabling fair comparison across models while providing a modular platform for integrating new methods, datasets, and encoders. A central contribution is the Clinical Report Quality Score (CRQS), a clinically grounded metric for evaluating factual correctness. CRQS maps reference and generated reports into structured clinical attributes and measures four complementary dimensions: clinical fact coverage, key information recall, hallucination rate, and clinical discordance, producing both an overall score and interpretable sub-scores. Experiments demonstrate that conventional language-generation metrics are weakly aligned with clinical correctness and frequently overestimate report quality. In contrast, CRQS reveals clinically meaningful differences between models and encoders that lexical metrics fail to capture. Together, the benchmark, public plug-and-play framework, and CRQS establish a reproducible foundation for rigorous evaluation of pathology report generation.