arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23794cs.CVcs.AI

PathScale-R1:用于病理图像分析的跨尺度推理

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对病理诊断多尺度需求及现有模型局限,引入抗捷径跨尺度病理推理框架,设计相关策略构建PathScale-VQA基准,经优化得到PathScale-R1,实验证明其在跨尺度推理任务性能及对单尺度病理VQA的有效迁移。

中文摘要 AI 辅助

病理诊断本质上是多尺度的,需要整合低倍镜下的整体组织结构和高倍镜下的细胞形态。然而,现有的病理基准和视觉语言模型大多在单尺度设置下开发,限制了多倍镜推理能力。此外,简单构建的视觉问答任务可能易受文本或表面视觉捷径影响。为解决这些局限,我们引入了抗捷径跨尺度病理推理的基准和训练框架。设计了语义推理问题的对抗性纯文本筛选策略和视觉定位问题的结构控制干扰采样策略。基于此构建了PathScale-VQA基准。在此基础上,PathScale-R1通过难度驱动的推理蒸馏监督微调及强化学习优化,实验证明其在跨尺度推理任务上的性能及向传统单尺度病理VQA的有效迁移。

英文摘要

Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at https://github.com/iMVR-PL/PathScale-R1.

发表机构

  • National University of Singapore(新加坡国立大学)
  • PuzzleLogic Pte Ltd(拼图逻辑私人有限公司)
  • Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院)

机构由 AI 辅助整理,请以论文原文为准。

↑