arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于前列腺组织病理学的病理学家注意力对齐报告生成

Pathologist Attention-Aligned Report Generation for Prostate Histopathology

Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty, Pierre Marza, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Paul Friedman, Bharat Ramlal, Beatrice Knudsen, Rajarsi Gupta, Joel Saltz, Prateek Prasanna, Gregory Zelinsky, Dimitris Samaras

arXiv 2607.19624首次发表:更新:

发表机构

Department of Computer Science, Stony Brook University; Department of Biomedical Informatics, Stony Brook University; Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit; Université Paris-Saclay, CentraleSupélec, MICS Laboratory; Department of Pathology and Laboratory Medicine, Northwell Health Laboratories; Department of Pathology, University of Utah School of Medicine; Department of Psychology, Stony Brook University(纽约州立大学石溪分校计算机科学系; 纽约州立大学石溪分校生物医学信息学系; 巴黎萨克雷大学、中央理工高等电力学院、古斯塔夫·鲁西研究所、法国国家健康与医学研究院、PRISM综合大学医院、癌症数据科学单元; 巴黎萨克雷大学、中央理工高等电力学院、MICS实验室; 诺斯韦尔健康实验室病理与检验医学部; 犹他大学医学院病理系; 纽约州立大学石溪分校心理学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究将人类注意力引入前列腺病理报告生成模型训练,收集多模态数据集,通过注意力对齐损失微调模型,在两个模型上评估,在报告生成和视觉问答任务中取得较好效果,模型注意力图与病理学家注意力更对齐。

AI 中文摘要

病理学家在癌症诊断过程中的视觉注意力分配是一个高度选择性的过程,对从全切片图像(WSIs)中提取的信息起着关键作用。人类注意力有助于诸如分类和分割等医学成像任务,并成为识别用于报告生成的诊断信息区域的强大语义线索。本文将人类注意力引入病理学家报告生成模型的训练中。为此,收集了一个包含121个前列腺WSIs的多模态人类注意力数据集,这些WSIs标注有病理学家与言语描述和光标移动同步的多尺度视口轨迹,涉及五个临床相关组件(如Gleason模式)。利用该数据集,通过注意力对齐损失对两个报告生成模型进行微调,该损失使模型对图像块的注意力正则化,以匹配病理学家注意力的分布。使用具有不同内部注意力机制的两个模型对前列腺癌报告生成和视觉问答方法进行评估。实验表明,在基于NLP的指标上平均提高了10.9%,在五个临床相关报告组件的准确性上提高了19.3%。此外,在推理时提取的模型注意力图与病理学家注意力更紧密对齐,通过突出影响输出最大的区域,为生成的报告提供了更强的视觉支持。

英文摘要

The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracted from whole-slide images (WSIs). Human attention helps medical imaging tasks such as classification and segmentation, and becomes a strong semantic cue for identifying diagnostically informative regions for report generation. In this paper, we introduce human attention into the training of pathologist report generation models. To this end, we collected a multimodal human-attention dataset of 121 prostate WSIs annotated with pathologists' multi-scale viewport trajectories synchronized with the pathologists' verbal descriptions and cursor movements for five clinically relevant components (e.g., Gleason patterns). Using this dataset, we finetune two report generation models with an attention-alignment loss that regularizes the model attention over image patches to match the distribution of pathologist attention. We evaluate our approach on prostate cancer report generation and visual question answering using two models with different internal attention mechanisms (i.e., how image tokens are integrated into the language decoder). Experiments show average gains of 10.9% on NLP-based metrics and 19.3% in accuracy across five clinically relevant report components. Further, model attention maps extracted at inference time, with minimal computational overhead, align more closely with pathologist attention, providing stronger visual support for the generated reports by highlighting the regions that most influence the output.

Comments11 pages, 4 figures, accepted for publication at the 29th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑