arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17926cs.CV

PerFact:用于3D脑部MRI报告生成的感知衍生事实提示方法

PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

Jianyu Sun, Zhenxuan Zhang, Guang Yang, Peter J. Lally

AI总结:

该研究针对3D脑部MRI报告生成,提出PerFact方法,以上游3D分割分类输出的结构化事实提示LoRA适配的视觉语言模型,证实grounding信息是报告质量的主导因素而非模型选择。

AI中文摘要:

放射科报告生成几乎完全在2D胸部X光片上成熟,提升报告质量的默认路径是更大的骨干网络或在医疗数据上预训练的骨干网络。我们在3D多序列脑部MRI(一种体积多疾病场景)上重新审视该假设,发现模型并非关键因素。零样本医疗和放射学视觉语言模型向脑部MRI的迁移效果很差,胸部X光片专家的表现最差;在三个模型家族中以相同方式微调的五个骨干网络,其规模相差一个数量级,性能差异却很小。报告质量的决定因素是注入提示的信息。我们将感知任务交给上游3D分割与分类,将其输出序列化为结构化事实语句,并用该语句提示经LoRA适配的视觉语言模型,我们将此方法命名为PerFact。在控制研究中,我们固定骨干网络、数据划分、目标报告和适配方式,仅改变注入的 grounding 信息,结果显示感知衍生事实优于检索到的既往报告;一旦存在事实,检索就变得多余;端到端预测的事实在推理时无需任何真实标注仍有效。预测事实与最优事实之间的剩余差距由事实的粒度而非生成器导致。封闭型视觉问答对报告质量无明显影响,尽管 grounding 源对其影响很小。在3D脑部MRI上,可控制的主导因素是 grounding 信息而非模型选择,这决定了报告质量。

英文摘要:

Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D multi-sequence brain MRI, a volumetric multi-disease regime, and find that the model is not the lever. Zero-shot medical and radiology vision-language models transfer poorly to brain MRI, with chest radiograph specialists failing most conspicuously, and five backbones fine-tuned identically across three model families and an order of magnitude in scale differ only marginally. What determines the quality of the report is the information injected into the prompt. We delegate perception to upstream 3D segmentation and classification, serialize their outputs into a structured fact sentence, and prompt a LoRA-adapted vision-language model with it; we call this \textbf{PerFact}. In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth annotation at inference. The residual gap between predicted and oracle facts is explained by the granularity of the facts rather than by the generator. Closed-ended visual question answering comes at no measurable cost to report quality, though the grounding source has little effect on it. On 3D brain MRI, grounding information, not model choice, is the dominant controllable factor in report quality.

↑