MIRA:面向智能体诊断的医学图像反射
MIRA: Medical Image Reflection for Agentic Diagnosis
浏览论文内容
中文总结 AI 辅助
该研究提出医学视觉诊断框架MIRA,通过两阶段训练优化工具使用与证据验证,在9个医学视觉推理基准上较Qwen3-VL-8B backbone提升7.44分,改善工具使用判断并减少有害判断。
中文摘要 AI 辅助
医学视觉智能体可使用工具检查图像、检索外部知识,但不加区分的工具使用可能引入嘈杂或误导性证据。因此,可靠的诊断不仅需要获取额外观测结果,还需验证工具操作是否必要、所得证据是否支持当前假设。我们提出MIRA(Medical Image Reflection for Agentic Diagnosis),这是一款用于自主证据搜索与反射验证的医学视觉诊断框架。MIRA会动态调用图像处理操作,包括缩放、定位、指向、旋转、测量,以及网络搜索,同时评估所获证据的相关性与一致性。我们通过两阶段训练策略开发MIRA:第一阶段,工具增强的蒙特卡洛树搜索数据引擎探索多样诊断假设,联合验证视觉定位准确率与语义一致性,以构建监督微调轨迹;第二阶段,强化学习通过在线反射原则进化进一步优化决策:将失败案例提炼为候选原则,仅保留能提升保留样本rollout奖励的原则。在9个医学视觉推理基准上,MIRA取得平均得分64.73,较其Qwen3-VL-8B backbone提升7.44分;还将有效工具使用判断从56.2%提升至73.8%,有害判断从8.9%降至1.6%。定性分析显示,MIRA可重新审查证据、修正过早结论并调整其工具使用策略。项目页面:this https URL
英文摘要
Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/
发表机构
- Tongji University(同济大学)
- Nanjing University(南京大学)
- Northwestern Polytechnical University(西北工业大学)
- People's Public Security University of China(中国人民公安大学)
机构由 AI 辅助整理,请以论文原文为准。