arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

观察、测量与推理:学习病理学中的视觉接地推理

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Xinyu Liu, Jiaming Yang, Jie Chen, Zhang Zhang, Yuhao Yi, Hong Bu, Jiancheng Lv

arXiv 2609.34277首次发表:更新:

发表机构

Sichuan University; West China Hospital, Sichuan University; National University of Singapore; Sun Yat-sen University(四川大学; 四川大学华西医院; 新加坡国立大学; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ASPECT方法,通过显式监督细胞外观与丰度、三阶段微调和强化学习,提升病理图像视觉接地推理,在PathoVernier基准上相对最强基线准确率提升19.2%。

AI 中文摘要

病理学评估依赖于识别组织学图像中细微的视觉细节。视觉语言模型(VLMs)越来越多地支持病理学解释,但其感知这些细节的能力仍然不足。这一弱点导致细胞观察不准确,即使最终答案正确,这种错误也可能持续存在。在本文中,我们提出ASPECT,通过显式监督细胞外观和丰度来改进视觉接地推理。ASPECT通过病理特征重建、细胞特征对齐和计数监督来训练中间视觉标记。三阶段监督微调教导模型感知、生成视觉标记并进行推理,随后进行强化学习,奖励答案正确性和与报告测量的一致性。我们还引入了PathoVernier,一个包含来自五个病理学数据集、涵盖四种细胞组成任务的759个专家审核问题的基准。它同时评估最终答案和中间测量,以揭示被答案准确性掩盖的错误。在PathoVernier上,ASPECT相对于最强基线Gemini-3.1-Pro取得了约19.2%的相对准确率提升,相对于其Qwen3-VL-8B骨干网络提升了99.3%,同时将衡量正确响应中计数错误的RAWR分别降低了28.1%和42.7%。ASPECT还在三个外部病理学基准上优于其骨干网络,这些基准涵盖分类和问答任务,超出了细胞组成任务的范围。

英文摘要

Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑