发表机构
Ulm University; Ulm University Hospital; Technical University of Munich (TUM); TUM University Hospital; Department of Radiation Oncology, TUM University Hospital(乌尔姆大学; 乌尔姆大学医院; 慕尼黑工业大学(TUM); 慕尼黑工业大学医院; 慕尼黑工业大学医院放射肿瘤科)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对医学视觉-语言模型空间推理薄弱的问题,提出模块化医学影像智能体,通过分阶段处理实现CT扫描空间关系验证,性能优于端到端基线,可作为未来医学影像智能体的构建块。
AI 中文摘要
可靠的空间理解是未来旨在支持放射报告生成和结构化图像理解的医学视觉-语言系统的重要前提。尽管现代视觉-语言模型(VLMs)在许多医学影像任务上表现出良好性能,但近期证据表明,它们在受控空间推理方面仍然薄弱,且往往无法可靠地将空间关系与图像证据关联。由于放射推理依赖于对解剖结构和发现的相对位置的理解,这种空间弱点对诊断准确性构成风险。我们提出一种用于轴向CT切片中二值空间关系验证的模块化医学影像智能体。该系统并非直接端到端预测空间答案,而是将任务分解为明确的阶段:语言解析、解剖定位和确定性几何验证。自然语言查询被转换为结构化关系元组,使用基于YOLO的检测器定位被查询器官,最终空间决策通过确定性几何规则从物体中心计算得出。我们在保留的MIRP空间QA基准上评估该方法,并将其与代表性的端到端VLM基线进行比较。表现最佳的混合配置达到94.1%的准确率和94.2%的F1值,在准确率上比直接Qwen2-VL提示高出42.5个百分点,同时保留可解释的中间表示和可审计的推理阶段。结果表明,显式模块化空间验证可作为未来面向报告的医学影像智能体的有前景的构建块。
英文摘要
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language models (VLMs) show promising performance on many medical imaging tasks, recent evidence suggests they remain weak in controlled spatial reasoning and often fail to reliably ground spatial relations in image evidence. Given that radiological reasoning hinges on understanding the relative positions of anatomical structures and findings, this spatial weakness poses risks to diagnostic accuracy. We present a modular medical imaging agent for binary spatial relation verification in axial CT slices. Instead of directly predicting spatial answers end-to-end, the system decomposes the task into explicit stages: language parsing, anatomical localization, and deterministic geometric verification. Natural-language queries are converted into structured relation tuples, queried organs are localized with a YOLO-based detector, and the final spatial decision is computed from object centers using deterministic geometric rules. We evaluate the approach on the held-out MIRP spatial QA benchmark and compare it against representative end-to-end VLM baselines. The best-performing hybrid configuration reaches 94.1% accuracy and 94.2% F1, outperforming direct Qwen2-VL prompting by 42.5 percentage points in accuracy, while preserving interpretable intermediate representations and auditable reasoning stages. The results suggest that explicit modular spatial verification can serve as a promising building block for future report-oriented medical imaging agents.
Journal refPublished at the MICCAI 2026 Agentic AI for Medicine Workshop