arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08713cs.CVcs.AI

分辨率与压缩结合:面向3D放射报告生成的高效视觉上下文

Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

  • German Cancer Research Center (DKFZ)(德国癌症研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jer… 展开作者

Jonathan Suprijadi, Raphael Stock, Moritz Langenberg, David Zimmerer, Kim-Celine Kahl, Stefan Denner, Yannick Kirchhoff, Karol Gotkowski, Maximilian Rokuss, Jeremias Traub, Tassilo Wald, Constantin Ulrich, Klaus Maier-Hein

AI总结:

该研究针对3D放射报告生成中视觉序列的计算瓶颈,通过实验评估不同视觉编码器、投影器和LLM,提出解剖学引导的ROI裁剪等策略,在两个数据集上取得最先进的临床macro F1结果。

AI中文摘要:

视觉-语言模型为自动化放射报告生成提供了有前景的路径,但将其应用于完整3D CT体积会带来巨大的计算挑战。现代基础视觉编码器(VE)每次扫描可产生数万个视觉 token,传递给大型语言模型(LLM)的视觉序列成为主要计算瓶颈。视觉-语言投影器可压缩该序列以减少计算,但可能丢弃临床相关细节;反之,有效的压缩可在保持下游 token 数量固定的同时容纳更高分辨率的输入。因此,如何在输入视野、空间分辨率和视觉-语言投影之间分配视觉 token 预算仍是一个未解决的设计问题。我们在两个大规模CT报告数据集(CT-RATE和Merlin)上系统评估了四种异构VE(基于CNN和ViT)、五种最高压缩比达64倍的 token 减少投影器及非减少的MLP投影器基线,以及五种指令微调的LLM(17亿至40亿参数)。在匹配的LLM token 预算下,解剖学引导的感兴趣区域裁剪是最一致的策略,在20种设置中的19种中提升了临床 macro F1值,对于3D ViT Primus编码器平均提升3.7个百分点,对于基于切片的2D ViT Curia编码器平均提升1.1个百分点。进一步提高输入分辨率高度依赖于投影器:PerceiverResampler与更高分辨率的Curia特征配对,在两个数据集的分辨率研究中产生最强配置。我们的最佳配置在测试集上达到了最先进的临床 macro F1值,在CT-RATE上达到49.5,在Merlin上达到49.0。代码和模型将在发表后发布。

英文摘要:

Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of vision tokens per scan, making the visual sequence passed to the large language model (LLM) a primary computational bottleneck. Vision-to-language projectors can compress this sequence to reduce computation, but may discard clinically relevant detail; conversely, effective compression can accommodate higher-resolution inputs while keeping the downstream token count fixed. How this vision-token budget should be allocated across input field of view, spatial resolution, and vision-to-language projection therefore remains an open design question. We systematically evaluate four heterogeneous VEs (CNN- and ViT-based), five token-reducing projectors at up to 64x compression alongside a non-reducing MLP projector baseline, and five instruction-tuned LLMs (1.7B--4B) on two large-scale CT report datasets (CT-RATE and Merlin). At matched LLM token budgets, anatomy-guided region of interest cropping is the most consistent strategy, improving clinical macro F1 in 19 of 20 settings by +3.7 points on average for the 3D ViT Primus encoder and +1.1 for the slice-based 2D ViT Curia encoder. Increasing input resolution further is strongly projector-dependent: the PerceiverResampler, paired with higher-resolution Curia features, yields the strongest configuration in the resolution study on both datasets. Our best configurations achieve state-of-the-art clinical macro F1 on the test sets, reaching 49.5 on CT-RATE and 49.0 on Merlin. Code and models will be published upon publication.

↑