arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15353cs.CV

用图约束的多实例学习工作流分解全切片图像报告生成

Decomposing Whole Slide Image Report Generation with Graph-Constrained Multiple Instance Learning Workflows

Antony Gitau, Martyna Borak, Bjørn-Jostein Singstad, Martin Paulson, Karl Thomas Hjelmervik, Ola Marius Lysaker, Veralia Gabriela Sanchez

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出图约束的多实例学习分解框架,结合Virchow2嵌入与器官条件图,提升WSI报告生成性能,明确器官路由是域转移瓶颈,支持错误定位。

中文摘要 AI 辅助

全切片图像(WSI)报告生成需要识别空间分布的病理特征并将其组织成连贯的诊断叙述。尽管直接的视觉到文本模型可以生成流畅的报告,但它们掩盖了视觉识别、结构化推理和语言生成的贡献及失败模式。我们提出了一个分解框架,其中冻结的Virchow2图像块嵌入通过多实例学习(MIL)分类头聚合,这些分类头回答器官特异性诊断问题。器官条件图约束这些答案组装成结构化推理链,语言模型将其实现为病理报告。在REG2026保留的2028张切片数据集上,所提出的工作流的链Jaccard分数达到0.702。若没有基于图的链构建,性能降至0.420;若器官特异性图被单个器官无关图替换,性能降至0.398;若语言模型从MIL预测自由构建链,性能降至0.371。使用相同的报告生成器,图结构化链将报告分数从0.330提高到0.495。在未微调的350张涵盖REG2026七个器官的外部TCGA WSI上,预期器官图在64.0%的案例中被选中,在86.6%的案例中排名前三。提供正确的器官图使与TCGA粗略原发诊断标签的一致性从61.8%提高到92.6%,确定器官路由是域转移下的主要瓶颈。总体而言,器官条件、图约束的链组装可改善结构化推理和报告生成,同时支持阶段特异性错误定位。

英文摘要

Whole-slide image (WSI) report generation requires recognizing spatially distributed pathological features and organizing them into a coherent diagnostic narrative. Although direct vision-to-text models can yield fluent reports, they obscure the contributions and failure modes of visual recognition, structured reasoning, and language generation. We propose a decomposed framework in which frozen Virchow2 tile embeddings are aggregated by multiple-instance learning (MIL) classification heads that answer organ-specific diagnostic questions. An organ-conditioned graph constrains the assembly of these answers into a structured reasoning chain, which a language model realizes as a pathology report. On the REG2026 held-out set of 2,028 slides, the proposed workflow achieved a chain-Jaccard score of 0.702. Performance fell to 0.420 without graph-based chain construction, 0.398 when the organ-specific graphs were replaced by a single organ-agnostic graph, and 0.371 when the language model constructed the chain freely from MIL predictions. Using the same report generator, graph-structured chains improved the report score from 0.330 to 0.495. On 350 external TCGA WSIs spanning the seven REG organs without fine-tuning, the expected organ graph was selected in 64.0% of cases and ranked among the top three in 86.6%. Providing the correct organ graph increased agreement with coarse TCGA primary-diagnosis labels from 61.8% to 92.6%, identifying organ routing as a main bottleneck under domain shift. Overall, organ-conditioned, graph-constrained chain assembly improves structured reasoning and report generation while enabling stage-specific error localization.

发表机构

  • Faculty of Technology, Natural Sciences and Maritime Sciences, University of South-Eastern Norway(东南挪威大学技术、自然科学与海洋科学学院)
  • Center for Cancer and Blood Diseases, Vestfold Hospital Trust(西福尔信托医院癌症与血液疾病中心)
  • Faculty of Biomedical Engineering, Silesian University of Technology(西里西亚工业大学生物医学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑