arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2506.11166cs.CVcs.AI

面向零样本诊断的测试时缩放与视觉-语言推理

Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning

  • Johns Hopkins University(约翰霍普金斯大学)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Ji Young Byun, Young-Jin Park, Navid Azizan, Rama Chellappa

更新

AI总结:

针对医学影像诊断中LLM推理应用不足、监督微调数据成本高的问题,提出基于测试时缩放的零样本框架,通过视觉-语言模型生成多解读再由LLM整合,在多模态医学影像上提升了诊断准确率与可靠性。

AI中文摘要:

作为患者护理的基石,临床决策对患者预后有显著影响,且可通过大语言模型(LLM)得到提升。尽管LLM已展现出卓越性能,但其在医学影像视觉问答中的应用,尤其是基于推理的诊断,仍有待深入探索。此外,由于数据有限且标注成本高昂,针对推理任务的监督微调基本不具备可行性。本研究提出了一种用于可靠医学影像诊断的零样本框架,通过测试时缩放增强LLM在临床场景中的推理能力。给定一张医学影像和一段文本提示,视觉-语言模型会处理该医学影像及对应的文本提示,生成视觉特征的多个描述或解读。这些解读随后被输入LLM,由测试时缩放策略将多个候选输出整合为可靠的最终诊断。我们在多种医学影像模态(包括放射学、眼科学和组织病理学)上评估了该方法,结果表明所提出的测试时缩放策略提升了我们的方法及基线方法的诊断准确率。此外,我们的实证分析显示,该方法允许在第一阶段使用无偏提示,可提升LLM生成诊断的可靠性并提高分类准确率。

英文摘要:

As a cornerstone of patient care, clinical decision-making significantly influences patient outcomes and can be enhanced by large language models (LLMs). Although LLMs have demonstrated remarkable performance, their application to visual question answering in medical imaging, particularly for reasoning-based diagnosis, remains largely unexplored. Furthermore, supervised fine-tuning for reasoning tasks is largely impractical due to limited data availability and high annotation costs. In this work, we introduce a zero-shot framework for reliable medical image diagnosis that enhances the reasoning capabilities of LLMs in clinical settings through test-time scaling. Given a medical image and a textual prompt, a vision-language model processes a medical image along with a corresponding textual prompt to generate multiple descriptions or interpretations of visual features. These interpretations are then fed to an LLM, where a test-time scaling strategy consolidates multiple candidate outputs into a reliable final diagnosis. We evaluate our approach across various medical imaging modalities -- including radiology, ophthalmology, and histopathology -- and demonstrate that the proposed test-time scaling strategy enhances diagnostic accuracy for both our and baseline methods. Additionally, we provide an empirical analysis showing that the proposed approach, which allows unbiased prompting in the first stage, improves the reliability of LLM-generated diagnoses and enhances classification accuracy.

↑