arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SlideBank:用于一致全切片推理的持久分层证据库

SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning

Beidi Zhao, Gexin Huang, Ciro Zhang, Anqi Li, Yusheng Tan, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li

arXiv 2609.00342首次发表:更新:

发表机构

University of British Columbia; Vector Institute; Harvard University; Rice University; University of Chicago; BC Cancer Agency(不列颠哥伦比亚大学; 矢量研究院; 哈佛大学; 莱斯大学; 芝加哥大学; 不列颠哥伦比亚癌症机构)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全切片图像推理的挑战,提出无需训练的SlideBank框架,通过分层证据库整合多尺度信息,在两个病理数据集上取得优异性能,提升了推理一致性并降低了成本。

AI 中文摘要

全切片图像(WSI)对视觉-语言推理具有挑战性,因为诊断相关的形态学特征稀疏、异质,且分布在千兆像素级图像和多个空间分辨率中。现有的WSI模型和病理智能体可聚合切片特征或主动获取证据,但探索后保留的信息往往难以在语义上访问,同时保持其与原始视觉证据的关联。我们引入SlideBank,这是一个无需训练的框架,将每个WSI表示为持久、概念索引且空间定位的证据库。SlideBank执行与问题无关的由粗到细探索,以识别信息丰富的区域和多尺度视图,将其转换为显式形态学观测,并将病理信号定位到其支持的切片和WSI坐标。在推理时,问题被路由到相关信号和证据尺度,通过基于置信度的跨级别共识整合关联的全局、区域和切片证据。在WSI-VQA和SlideBench-BCNB上的实验表明,使用Patho-R1时,SlideBank在WSI-VQA上达到52.77%的准确率,使用Quilt-LLaVA时,在SlideBench-BCNB上达到50.92%的平均准确率,而结构化信号引导的检索始终优于随机证据采样。在重复查询中重用同一库进一步实现了超过99%的复述一致性,并通过持久证据重用大幅降低了摊销推理成本。

英文摘要

Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggregate slide features or actively acquire evidence, but the information retained after exploration is often difficult to access semantically while preserving its connection to the original visual evidence. We introduce SlideBank, a training-free framework that represents each WSI as a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank performs question-independent coarse-to-fine exploration to identify informative regions and multi-scale views, converts them into explicit morphological observations, and grounds pathology signals to their supporting patches and WSI coordinates. At inference time, questions are routed to relevant signals and evidence scales, and the linked global, regional, and patch evidence is integrated through confidence-based cross-level consensus. Experiments on WSI-VQA and SlideBench-BCNB show that with Patho-R1, SlideBank reaches 52.77% on WSI-VQA and with Quilt-LLaVA, it reaches 50.92% average accuracy on SlideBench-BCNB, while structured signal-guided retrieval consistently outperforms random evidence sampling. Reusing the same bank across repeated queries further achieves over 99% rephrasing consistency and substantially reduces amortized inference cost through persistent evidence reuse.

Comments23 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑