arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越流畅性:脊柱MRI报告生成的临床基准与异常增强基线

Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation

Bruno Palau, Franziska Vogt, Daria Laslo, Haobo Li, Ender Konukoglu, Maria Monzon, Catherine R. Jutzeler

arXiv 2608.07117首次发表:更新:

发表机构

ETH Zurich; Swiss Institute of Bioinformatics (SIB)(苏黎世联邦理工学院; 瑞士生物信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对放射学报告的耗时与阅片差异问题,构建腰椎MRI的VLM临床基准,提出半监督U-Net++生成异常热图增强VLM的框架,提升诊断可靠性与可解释性。

AI 中文摘要

放射学报告耗时且存在阅片者间差异,使得自动报告生成成为视觉语言模型(VLMs)的极具吸引力的临床应用。我们在腰椎MRI上对最先进的VLMs进行基准测试,重点关注诊断准确性,并证明标准的词汇和语义指标无法很好地反映临床正确性:流畅、结构良好的报告可能得分很高,同时包含有临床意义的诊断错误。为解决这一失效模式,我们提出一种与架构无关的框架,该框架通过半监督U-Net++模型生成的空间定位的椎间盘水平异常热图来增强VLM的输入。这些热图既通过明确的视觉定位提高了解剖学敏感性,又为临床监督提供了独立的可解释性输出,使我们更接近用于腰椎MRI解读的诊断可靠、视觉定位的VLMs。

英文摘要

Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.

CommentsMaria Monzon and Catherine R. Jutzeler contributed equally as shared last authors. Accepted at the CV4Clinic Workshop, CVPR 2026

Journal refProc. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 6759-6770

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑