arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39566cs.CV

从给定证据到主动收集证据:面向纵向医学推理的智能体学习

From Given to Gathered Evidence: Agentic Learning for Longitudinal Medical Reasoning

Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出CASE智能体学习框架,通过工具使用和强化学习训练紧凑视觉-语言模型,在UK Biobank纵向多模态基准上主动寻求证据,提升医学推理准确性,相对GPT-5.4和Claude Opus 4.8分别提高超16%和10%。

中文摘要 AI 辅助

基础模型可以通过工具使用框架作为临床智能体。然而,传统的医学基准评估的是对预先选择证据的推理能力,而非在临床记录和纵向影像中主动寻找证据的能力。我们提出CASE:一系列角色特定的临床证据寻求智能体,以及一个工具使用框架和一个用于紧凑视觉-语言策略模型的智能体后训练框架。我们进一步引入一个基于英国生物银行构建的纵向多模态基准,包含来自4,739名参与者的真实世界ICD-10编码诊断的50,401个临床问题。每个问题关联一个患者特定环境,包含临床背景和基线与随访访问的多序列MRI,智能体自主选择要检查和比较的访问、器官、模态、切片和专家工具。监督微调从14,734条前沿模型交互轨迹中迁移证据寻求工作流,随后在智能体自身环境交互上进行智能体强化学习。特权在线自蒸馏和基于规则的LLM反馈完善了从证据到结论的推理,而不规定工具序列。实验表明,CASE超越了问答模仿,转向可迁移的调查策略,加强了基于证据的纵向推理。在匹配的评估条件下,我们基于Qwen3-VL-8B的智能体在答案准确性上相对于GPT-5.4和Claude Opus 4.8分别实现了超过16%和10%的相对提升。代码将在该https URL上提供。

英文摘要

Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected evidence rather than the ability to seek it across clinical records and longitudinal imaging. We propose CASE: a series of role-specific Clinical Agents for Seeking Evidence, together with a tool-use harness and an agentic post-training framework for compact vision-language policy models. We further introduce a longitudinal multimodal benchmark built on UK Biobank, comprising 50,401 clinical questions derived from real-world ICD-10-coded diagnoses of 4,739 participants. Each question links to a patient-specific environment containing clinical context and multi-sequence MRI from baseline and follow-up visits, where agents autonomously select which visits, organs, modalities, slices, and specialist tools to inspect and compare. Supervised fine-tuning transfers evidence-seeking workflows from 14,734 frontier-model interaction trajectories, followed by agentic reinforcement learning on the learner's own environment interactions. Privileged on-policy self-distillation and rubric-based LLM feedback refine evidence-to-conclusion reasoning without prescribing tool sequences. Experiments show that CASE moves beyond question-answer imitation toward transferable investigation policies, strengthening evidence-grounded longitudinal reasoning. Under matched evaluation conditions, our Qwen3-VL-8B based agent achieves over 16% and 10% relative improvements in answer accuracy over GPT-5.4 and Claude Opus 4.8. Code will be available at https://github.com/VinyehShaw/CASE.

发表机构

  • Durham University(杜伦大学)
  • DAMO Academy, Alibaba Group(阿里巴巴达摩院)
  • Hupan Laboratory(湖畔实验室)
  • University of the Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

↑